UniParser/MolParser-7M
收藏资源简介:
MolParser-7M是一个用于分子结构视觉识别的数据集,包含近800万对图像-SMILES数据。图像标注采用扩展的SMILES格式。数据集包括以下子集:MolParser-7M(预训练)包含超过770万合成训练数据;MolParser-SFT包含人类标记的真实分子图像用于微调;MolParser-Val包含一个小型验证集,用于训练过程中快速验证模型能力;WildMol Benchmark包含从真实专利或论文中裁剪的2万分子结构图像,分为test_simple_10k(WildMol-10k)和test_markush_10k(WildMol-10k-M)两个子集。该数据集基于论文《MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild》(ICCV2025接受)提出,主要用于非商业用途。
MolParser-7M is a dataset for visual recognition of molecule structures, containing nearly 8 million paired image-SMILES data. The image captions are in an extended-SMILES format. The dataset includes the following subsets: MolParser-7M (Pretrain) with over 7.7M synthetic training data; MolParser-SFT with human-labeled real molecule figures for fine-tuning; MolParser-Val, a small validation set for quick model validation during training; and WildMol Benchmark with 20k molecule structure images cropped from real patents or papers, divided into test_simple_10k (WildMol-10k) and test_markush_10k (WildMol-10k-M) subsets. It is proposed in the paper MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild (ICCV2025 accept) and intended for non-commercial use only.




