遇见数据集

CSS and Benchmark Datasets of GeminiMol

收藏
Zenodo2024-03-06 更新2026-05-26 收录
官方服务:

资源简介:

The molecular representation model is a neural network that converts molecular representations (SMILES, Graph) into feature vectors, that carries the potential to be applied across a wide scope of drug discovery scenarios. However, current molecular representation models have been limited to 2D or static 3D structures, overlooking the dynamic nature of small molecules in solution and their ability to adopt flexible conformational changes crucial for drug-target interactions. To address this limitation, we propose a novel strategy that incorporates the conformational space profile into molecular representation learning. By capturing the intricate interplay between molecular structure and conformational space, our strategy enhances the representational capacity of our model named GeminiMol. Consequently, when pre-trained on a miniaturized molecular dataset, the GeminiMol model demonstrates a balanced and superior performance not only on traditional molecular property prediction tasks but also on zero-shot learning tasks, including virtual screening and target identification. By capturing the dynamic behavior of small molecules, our strategy paves the way for rapid exploration of chemical space, facilitating the transformation of drug design paradigms. In this study, a diverse collection of 39,290 molecules was employed for conformational searching and shape alignment to generate a comprehensive dataset of molecular conformational space similarity. To assess the model's performance, the benchmark datasets comprising over millions molecules was utilized for downstream tasks. Here, we provide all the training and benchmarking data used for this study to facilitate the reproducibility of the work.

分子表征模型(molecular representation model)是一类将分子表征形式(简化分子线性输入规范(SMILES)、分子图(Graph))转换为特征向量的神经网络,具备在广泛药物发现场景中应用的潜力。然而,当前的分子表征模型仅局限于二维或静态三维结构,忽略了溶液中小分子的动态特性,以及其为适配药物-靶点相互作用而发生柔性构象变化的关键能力。 为解决这一局限,我们提出了一种将构象空间特征纳入分子表征学习的全新策略。通过捕捉分子结构与构象空间之间的复杂相互作用,该策略提升了我们所构建的GeminiMol模型的表征能力。因此,在小型分子数据集上完成预训练后,GeminiMol模型不仅在传统分子性质预测任务中展现出均衡且优异的性能,还在包括虚拟筛选、靶点识别在内的零样本(zero-shot)学习任务中表现出色。通过捕捉小分子的动态行为,该策略为化学空间的快速探索铺平了道路,助力药物设计范式的革新。 本研究选取了包含39290个分子的多样化集合,用于构象搜索与形状对齐,以构建一套涵盖分子构象空间相似性的全面数据集。为评估模型性能,我们使用了包含数百万个分子的基准数据集来完成下游任务。本文公开了本研究中使用的全部训练与基准测试数据,以推动该研究成果的可复现性。

创建时间:
2023-12-14
二维码
社区交流群
二维码
科研交流群
商业服务