遇见数据集

MolCap

收藏
魔搭社区2026-07-09 更新2026-07-15 收录
官方服务:

资源简介:

# Molecule Image Caption Dataset 📘 Dataset Summary MolCap is a large-scale multi-modal molecular dataset with over 320k molecular images and detailed captions. The images are rendered by RDKit with random perturbations, and the captions, derived from PubChem descriptions, are cleaned and rewritten using GPT-4o. Each caption includes the canonical SMILES, the E-SMILES representation (introduced in “MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild“, ICCV2025), as well as structural details, physicochemical properties, and other relevant descriptors. 📖 Citation If you use this dataset in your research, please cite: ``` @article{fang2025uni, title={Uni-Parser Technical Report}, author={Fang, Xi and Tao, Haoyi and Yang, Shuwen and Zhong, Suyang and Lu, Haocheng and Lyu, Han and Huang, Chaozheng and Li, Xinyu and Zhang, Linfeng and Ke, Guolin}, journal={arXiv preprint arXiv:2512.15098}, year={2025} } ``` ``` @article{fang2024molparser, title={Molparser: End-to-end visual recognition of molecule structures in the wild}, author={Fang, Xi and Wang, Jiankun and Cai, Xiaochen and Chen, Shangqian and Yang, Shuwen and Tao, Haoyi and Wang, Nan and Yao, Lin and Zhang, Linfeng and Ke, Guolin}, journal={arXiv preprint arXiv:2411.11098}, year={2024} } ```

提供机构:
maas
创建时间:
2026-01-19
二维码
社区交流群
二维码
科研交流群
商业服务