MCMoD
收藏资源简介:
MCMoD数据集是由北京大学深圳研究生院开发的一个大规模多条件分子设计数据集,旨在解决化学语言模型在分子设计与生成中的语义鸿沟问题。该数据集包含超过100万条分子数据,涵盖了合成分子、天然产物和蛋白质配体等多种分子类型,并提供了文本描述、分子片段和化学性质等多重控制条件。数据来源包括PubChem、ZINC、ChEBI等知名数据库,并通过RDKit等工具进行标准化处理。MCMoD数据集不仅支持多目标联合控制,还创新性地利用分子片段序列作为文本与分子模态之间的桥梁,推动了化学语言模型在药物开发和化学工程等领域的实际应用。
The MCMoD dataset is a large-scale multi-condition molecular design dataset developed by Peking University Shenzhen Graduate School, aiming to address the semantic gap issue in molecular design and generation for chemical language models. This dataset contains over 1 million molecular entries, covering diverse molecular types such as synthetic molecules, natural products, and protein ligands, and provides multiple control conditions including textual descriptions, molecular fragments, and chemical properties. The data is sourced from well-known databases including PubChem, ZINC, ChEBI, and others, and has been standardized through tools such as RDKit. The MCMoD dataset not only supports multi-objective joint control, but also innovatively employs molecular fragment sequences as a bridge between text and molecular modalities, advancing the practical applications of chemical language models in fields such as drug development and chemical engineering.

- 1Navigating Chemical-Linguistic Sharing Space with Heterogeneous Molecular Encoding北京大学深圳研究生院电子与计算机工程学院 · 2024年



