遇见数据集

MDD-Molecular Dynamics Dataset: Collection of protein-ligand complex simulations

收藏
Zenodo2026-02-27 更新2026-05-26 收录
官方服务:

资源简介:

Dataset is part of the paper: https://chemrxiv.org/engage/chemrxiv/article-details/664c73f6418a5379b0de8152. This dataset consists of molecular dynamics (MD) simulations of 862 unique protein-ligand complexes, covering a wide range of protein families and diverse chemical classes of ligands. It is derived from publicly available repositories and represents the largest single source of MD simulations to date. All protein-ligand complexes included in the dataset were prepared following a standardized protocol. Missing atoms in the protein structures were added using the PDBFixer tool. The protein targets were parameterized using the AMBER99SB-ILDN force field, while ligands were parameterized with the ANTECHAMBER module within the ACPYPE tool. Ligand partial charges were determined to match the quantum-mechanically generated electrostatic potential via the Restrained Electrostatic Potential (RESP) method, and the remaining parameters were set using the GAFF2 force field. The molecular dynamics simulations were performed using GROMACS. The simulations were configured in a cubic simulation box with periodic boundary conditions and employed a TIP3P water model within an electrostatically neutral environment. The simulation protocol included an initial minimization cycle, followed by temperature equilibration in the NVT ensemble and pressure equilibration in the NPT ensemble. Production simulations were conducted over a period of 200 ns, with a timestep of 100 ps. Constructing a large, representative set of MD simulations poses challenges due to the high computational costs and complexities associated with preparing molecular systems. Moreover, given the limited number of suitable training examples (complexes) and the large volume of MD data from each simulation, careful filtering and feature selection are crucial. This dataset is valuable for exploring how molecular dynamics simulation data can be integrated with protein-ligand binding affinity prediction tasks, an essential component of in silico drug discovery pipelines. MD simulations, in particular, offer a dynamic view by illustrating the temporal interactions within protein-ligand complexes, potentially providing additional insights for affinity and specificity estimates.

本数据集隶属于以下论文:https://chemrxiv.org/engage/chemrxiv/article-details/664c73f6418a5379b0de8152。 本数据集包含862个独特的蛋白质-配体复合物的分子动力学(Molecular Dynamics, MD)模拟数据,覆盖了众多蛋白质家族以及种类多样的配体化学类别。该数据集源自公开可用的数据库,是目前为止规模最大的单一来源MD模拟数据集。 本数据集收录的所有蛋白质-配体复合物均按照标准化流程完成制备。蛋白质结构中缺失的原子通过PDBFixer工具进行补全;蛋白质靶点采用AMBER99SB-ILDN力场进行参数化,配体则通过ACPYPE工具内置的ANTECHAMBER模块完成参数化。配体的部分电荷通过约束静电势(Restrained Electrostatic Potential, RESP)方法计算得到,以匹配量子力学生成的静电势,其余参数则采用GAFF2力场进行设置。所有分子动力学模拟均通过GROMACS软件完成,模拟体系设置为立方体模拟盒,采用周期性边界条件,并在静电中性环境中使用TIP3P水模型。模拟流程包含初始能量最小化循环,随后分别在NVT系综中进行温度平衡,以及在NPT系综中进行压力平衡;生产模拟时长为200纳秒,时间步长为100皮秒。 构建大规模且具有代表性的MD模拟数据集面临诸多挑战,这源于分子体系制备过程中高昂的计算成本与复杂程度。此外,鉴于可用的合适训练样本(复合物)数量有限,且单次模拟产生的MD数据体量庞大,因此严格的筛选与特征选择环节至关重要。本数据集对于探索如何将分子动力学模拟数据与蛋白质-配体结合亲和力预测任务相结合具有重要价值,而该任务是in silico药物发现流程中的核心环节之一。尤为重要的是,MD模拟能够通过展示蛋白质-配体复合物内随时间变化的相互作用,提供动态视角,有望为结合亲和力与特异性评估提供额外的科学见解。

提供机构:
Zenodo
创建时间:
2024-05-16
二维码
社区交流群
二维码
科研交流群
商业服务