遇见数据集

docling-project/MarkushGrapher-Datasets

收藏
Hugging Face2025-06-05 更新2026-01-03 收录
官方服务:

资源简介:

--- license: cc-by-4.0 configs: - config_name: m2s data_files: - split: test path: "m2s/test/*.arrow" - config_name: uspto-markush data_files: - split: test path: "uspto-markush/test/*.arrow" - config_name: markushgrapher-synthetic data_files: - split: test path: "markushgrapher-synthetic/test/*.arrow" - config_name: markushgrapher-synthetic-training data_files: - split: train path: "markushgrapher-synthetic-training/train/*.arrow" - split: test path: "markushgrapher-synthetic-training/test/*.arrow" --- <div align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/64d38f55f8082bf19b7339e0/V43x-_idEdiCQIfbm0eVM.jpeg" alt="Description" width="800"> </div> This repository contains datasets introduced in [MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures](https://github.com/DS4SD/MarkushGrapher). Training: - **MarkushGrapher-Synthetic-Training**: This set contains synthetic Markush structures used for training MarkushGrapher. Samples are synthetically generated using the following steps: (1) SMILES to CXSMILES conversion using RDKit; (2) CXSMILES rendering using CDK; (3) text description generation using templates; and (4) text description augmentation with LLM. Benchmarks: - **M2S**: This set contains 103 real Markush structures from patent documents. Samples are crops of both Markush structure backbone images and their textual descriptions. They are extracted from documents published in USPTO, EPO and WIPO. - **USPTO-Markush**: This set contains 75 real Markush structure backbone images from patent documents. They are extracted from documents published in USPTO. - **MarkushGrapher-Synthetic**: This set contains 1000 synthetic Markush structures. Its images are sampled such that overall, each Markush features (R-groups, ’m’ and ’Sg’ sections) is represented evenly. An example of how to read the dataset is provided in [dataset_explorer.ipynb](https://huggingface.co/datasets/ds4sd/MarkushGrapher-Datasets/blob/main/dataset_explorer.ipynb).

许可证:CC BY 4.0 配置项: - 配置名称:m2s 数据文件: - 数据集拆分:测试集 路径:"m2s/test/*.arrow" - 配置名称:uspto-markush 数据文件: - 数据集拆分:测试集 路径:"uspto-markush/test/*.arrow" - 配置名称:markushgrapher-synthetic 数据文件: - 数据集拆分:测试集 路径:"markushgrapher-synthetic/test/*.arrow" - 配置名称:markushgrapher-synthetic-training 数据文件: - 数据集拆分:训练集 路径:"markushgrapher-synthetic-training/train/*.arrow" - 数据集拆分:测试集 路径:"markushgrapher-synthetic-training/test/*.arrow" --- <div align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/64d38f55f8082bf19b7339e0/V43x-_idEdiCQIfbm0eVM.jpeg" alt="数据集示意图" width="800"> </div> 本仓库包含来自论文《马克什结构识别器:马克什结构的视觉与文本联合识别》(MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures)的数据集,论文链接为https://github.com/DS4SD/MarkushGrapher。 训练集: - **马克什结构识别器合成训练集(MarkushGrapher-Synthetic-Training)**:该数据集包含用于训练马克什结构识别器的合成马克什结构样本。样本通过以下步骤合成生成:(1) 使用RDKit将SMILES转换为CXSMILES;(2) 使用CDK渲染CXSMILES图像;(3) 基于模板生成文本描述;(4) 使用大语言模型(Large Language Model,简称LLM)对文本描述进行增强。 基准测试集: - **M2S**:该数据集包含103个来自专利文献的真实马克什结构样本。样本均为马克什结构骨架图像及其对应文本描述的裁剪片段,提取自美国专利商标局(USPTO)、欧洲专利局(EPO)以及世界知识产权组织(WIPO)发布的专利文档。 - **USPTO-Markush**:该数据集包含75个来自专利文献的真实马克什结构骨架图像,均提取自美国专利商标局(USPTO)发布的专利文档。 - **马克什结构识别器合成基准测试集(MarkushGrapher-Synthetic)**:该数据集包含1000个合成马克什结构样本。其图像采样方式确保所有马克什结构特征(R基团、“m”与“Sg”区段)的分布均较为均匀。 本仓库提供了数据集读取示例,详见[dataset_explorer.ipynb](https://huggingface.co/datasets/ds4sd/MarkushGrapher-Datasets/blob/main/dataset_explorer.ipynb).

提供机构:
docling-project
二维码
社区交流群
二维码
科研交流群
商业服务