Lig-PCDB: Labeled Databases of X-ray Ligands Images in 3D Point Clouds and Validated Deep Learning Models
收藏资源简介:
LigPCDS: Labeled Dataset of X-ray Protein Ligand 3D Images in Point Clouds and Validated Deep Learning Models The difference electron density from X-ray protein crystallography was used to create the first dataset of labeled images of ligands in 3D point clouds, named LigPCDS. Four proposed vocabularies were validated by successfully training good performance deep learning models for the semantic segmentation of a stratified dataset from Lig-PCDB. The data from organic molecules (ligands) was obtained from the world Protein Data Bank with resolutions ranging from 1.5 to 2.2 Å. The ligands' images were interpolated from their calculated difference electron density map in a 3D grid-like bounding box, around their atomic positions, and stored in point clouds. A grid spacing of 0.5 Å gave the best results. The density value of the grid points was used as feature. The labeling approach used the structure of the ligands to propose vocabularies of chemical classes based on the chemical atoms themselves and their cyclic substructures. These annotations were applied pointwise to the ligands' images using an atomic sphere model. The databases and validated models may be used to tackle problems regarding known and unknown ligand building to drug discovery and fragment screening pipelines. The four validated deep learning models are: (i) the LigandRegion, composed by generic atoms of any type; (ii) the AtomCycle, composed by generic atoms outside cycles and generic cycles; (iii) the AtomC347CA56, composed by generic atoms outside cycles, not aromatic cycles of size 3 to 7 and aromatic cycles of size 5 and 6; and (iv) the AtomSymbolGroups, composed by the atoms symbols with groupings. The mean accuracy of these models in their cross-validation was between 49.7% and 77.4% in terms of Intersection over Union (mIoU) metric and between 62.4% and 87.0% in F1-score (mF1). The code used to create and validated the Lig-PCDB is available at the following repository: https://github.com/danielatrivella/np3_ligand This repository also contains the NP³ Blob Label application for ligand building using the validated deep learning models from Lig-PCDB. License LigPCDS by Cristina Freitas Bazzano, Luiz G. Alves,Guilherme P. Telles, Daniela B. B. Trivella is marked with CC0 1.0 Universal .
LigPCDS:基于点云的X射线蛋白质配体三维图像标注数据集与经过验证的深度学习模型 本研究利用X射线蛋白质晶体学(X-ray protein crystallography)得到的差分电子密度数据,构建了首个配体三维点云图像标注数据集,命名为LigPCDS。 研究团队通过对来自Lig-PCDB的分层数据集进行语义分割训练,成功得到性能优异的深度学习模型,以此验证了四种所提出的词汇表。本研究中的有机分子(配体)数据均取自全球蛋白质数据银行(Protein Data Bank, PDB),其分辨率范围为1.5至2.2 Å。配体图像通过其原子位置周围三维网格状边界框内的计算差分电子密度图插值生成,并以点云形式存储,其中0.5 Å的网格间距取得了最优效果。网格点的密度值被用作特征。标注方法以配体结构为基础,根据化学原子本身及其环状子结构提出化学类别词汇表,并通过原子球模型将上述标注逐点应用于配体图像。该数据集与经过验证的模型可用于解决药物发现与片段筛选流程中涉及的已知与未知配体构建相关问题。 本次验证的四种深度学习模型分别为:(i) 配体区域模型(LigandRegion):由任意类型的通用原子构成;(ii) 原子环模型(AtomCycle):由环外通用原子与通用环构成;(iii) AtomC347CA56模型:由环外、3至7元非芳香环以及5元和6元芳香环对应的通用原子构成;(iv) 原子符号分组模型(AtomSymbolGroups):基于原子符号进行分组构建。经交叉验证,这些模型的平均交并比(mean Intersection over Union, mIoU)指标得分介于49.7%至77.4%之间,平均F1分数(mean F1-score, mF1)得分介于62.4%至87.0%之间。 用于构建与验证Lig-PCDB的代码可通过以下仓库获取:https://github.com/danielatrivella/np3_ligand。该仓库同时包含用于基于Lig-PCDB的验证深度学习模型进行配体构建的NP³ Blob Label应用程序。 许可证 本数据集LigPCDS由Cristina Freitas Bazzano、Luiz G. Alves、Guilherme P. Telles、Daniela B. B. Trivella创建,采用CC0 1.0 Universal 通用公共领域贡献协议1.0进行授权。



