遇见数据集

Kinodata-3D: an in silico kinase-ligand complex dataset for kinase-focused machine learning.

收藏
Zenodo2024-03-22 更新2026-05-29 收录
官方服务:

资源简介:

Project Description Drug discovery pipelines nowadays rely on machine learning models to explore and evaluate large chemical spaces. While the inclusion of 3D complex information is considered to be beneficial, structural ML for affinity prediction suffers from data scarcity. We provide kinodata-3D, a dataset of ~138 000 docked complexes to enable more robust training of 3D-based ML models for kinase activity prediction (see github.com/volkamerlab/kinodata-3D-affinity-prediction). Dataset 1. Data This data set consists of three-dimensional protein-ligand complexes that were generated using computational docking from the OpenEye toolkit. The modeled proteins cover the kinase family for which a fair amount of structural data, i.e. co-crystallized protein-ligand complexes in the PDB, enriched through KLIFS annotations, is available. This enables us to use template docking (OpenEye’s POSIT functionality) in which the ligand placement is guided according to a similar co-crystallized ligand pose. The kinase-ligand pairs to dock are sourced from binding assay data via the public ChEMBL archive, version 33. In particular, we use kinase activity data as curated through the OpenKinome kinodata project. The final protein-ligand complexes are annotated with a predicted RMSD of the docked poses. The RMSD model is a simple neural network trained on a kinase-docking benchmark data set using ligand (fingerprint) similarity, docking score (ChemGauss 4), and Posit probability (see kinodata-3D repository). The final data set contains in total 138 286 deduplicated kinase-ligand pairs, covering ~98 000 distinct compounds and ~271 distinct kinase structures. 2. File structure The archive kinodata_3d.zip uses the following file structure data/raw | kinodata_docked_with_rmsd.sdf.gz | pocket_sequences.csv | mol2/pocket | 1_pocket.mol2 | ... The file kinodata_docked_with_rmsd.sdf.gz contains the docked ligand poses and the information on the protein-ligand pair inherited from kinodata. The protein pockets located in mol2/pocket are stored according to the MOL2 file format. The pocket structures were sourced from KLIFS (klifs.net) and complete the poses in the aforementioned SDF file. The files are named {klifs_structure_id}_pocket.mol2. The structure ID is given in the SDF file along with the ligand poses. The file pocket_sequences.csv contains all KLIFS pocket sequences relevant to the kinodata-3D dataset. 3. Related code The code used to create the poses can be found in the kinodata-3D repository. The docking pipeline makes heavy use of the kinoml framework, which in turn uses OpenEye's Posit template docking implementation. The details of the original pipeline can also be found in the manuscript by Schaller et al. (2023). Benchmarking Cross-Docking Strategies for Structure-Informed Machine Learning in Kinase Drug Discovery. bioRxiv.

项目概述 当前药物发现管线多依托机器学习模型探索与评估海量化学空间。尽管引入三维复合物信息被认为可有效提升模型性能,但用于亲和力预测的结构化机器学习模型仍面临数据匮乏的困境。本研究发布kinodata-3D数据集,该数据集包含约13.8万个对接复合物,可支撑基于三维结构的激酶活性预测机器学习模型的稳健训练(详见github.com/volkamerlab/kinodata-3D-affinity-prediction)。 数据集 1. 数据 本数据集包含通过OpenEye工具包执行计算对接得到的三维蛋白质-配体复合物。所建模的蛋白质覆盖激酶家族,该家族拥有较为丰富的结构数据,即通过KLIFS注释富集得到的蛋白质数据库(PDB)共结晶蛋白质-配体复合物。这使得我们可以采用模板对接(OpenEye的POSIT功能),即依据相似共结晶配体的构象指导配体的空间排布。待对接的激酶-配体对源自公开的ChEMBL数据库33版中的结合测定数据,具体而言,我们使用了由OpenKinome kinodata项目整理的激酶活性数据。最终的蛋白质-配体复合物均标注了对接构象的预测均方根偏差(RMSD)。该RMSD预测模型为简单神经网络,基于激酶对接基准数据集训练得到,输入特征包括配体(指纹)相似度、对接得分(ChemGauss 4)以及POSIT概率(详见kinodata-3D代码仓库)。 最终数据集共包含138286条去重后的激酶-配体对,覆盖约98000种不同化合物与271种不同激酶结构。 2. 文件结构 压缩包kinodata_3d.zip采用如下文件结构: data/raw | kinodata_docked_with_rmsd.sdf.gz | pocket_sequences.csv | mol2/pocket | 1_pocket.mol2 | ... 其中,`kinodata_docked_with_rmsd.sdf.gz` 文件存储了对接得到的配体构象,以及从kinodata数据集继承的蛋白质-配体对相关信息。`mol2/pocket/` 目录下的蛋白质结合口袋采用MOL2文件格式存储。结合口袋结构源自KLIFS数据库(klifs.net),用于补充前述结构数据文件(SDF)中的构象信息。口袋文件以`{klifs_structure_id}_pocket.mol2`的格式命名,SDF文件中会随配体构象一同提供对应的结构ID。`pocket_sequences.csv` 文件包含kinodata-3D数据集涉及的所有KLIFS口袋序列。 3. 相关代码 用于生成配体构象的代码可在kinodata-3D代码仓库中获取。本对接管线大量使用了kinoml框架,该框架依托OpenEye的POSIT模板对接实现。原始管线的详细信息可参见Schaller等人2023年发表的预印本手稿:《Benchmarking Cross-Docking Strategies for Structure-Informed Machine Learning in Kinase Drug Discovery》,发表于bioRxiv平台。

提供机构:
Zenodo
创建时间:
2023-12-21
二维码
社区交流群
二维码
科研交流群
商业服务