Reading PDB: Perception of Molecules from 3D Atomic Coordinates
收藏资源简介:
The analysis of small molecule crystal structures is a common way to gather valuable information for drug development. The necessary structural data is usually provided in specific file formats containing only element identities and three-dimensional atomic coordinates as reliable chemical information. Consequently, the automated perception of molecular structures from atomic coordinates has become a standard task in cheminformatics. The molecules generated by such methods must be both chemically valid and reasonable to provide a reliable basis for subsequent calculations. This can be a difficult task since the provided coordinates may deviate from ideal molecular geometries due to experimental uncertainties or low resolution. Additionally, the quality of the input data often differs significantly thus making it difficult to distinguish between actual structural features and mere geometric distortions. We present a method for the generation of molecular structures from atomic coordinates based on the recently published NAOMI model. By making use of this consistent chemical description, our method is able to generate reliable results even with input data of low quality. Molecules from 363 Protein Data Bank (PDB) entries could be perceived with a success rate of 98%, a result which could not be achieved with previously described methods. The robustness of our approach has been assessed by processing all small molecules from the PDB and comparing them to reference structures. The complete data set can be processed in less than 3 min, thus showing that our approach is suitable for large scale applications.
小分子晶体结构分析是获取药物研发宝贵信息的常用途径。所需的结构数据通常以特定文件格式提供,此类文件仅包含元素种类与三维原子坐标这类可靠的化学信息。因此,从原子坐标出发实现分子结构的自动化识别,已成为化学信息学领域的一项标准任务。通过此类方法生成的分子需同时满足化学合法性与合理性,方能为后续计算提供可靠基础。然而该任务颇具挑战:由于实验不确定性或分辨率偏低,提供的原子坐标可能偏离理想的分子几何构型;此外,输入数据的质量往往差异显著,这使得区分真实结构特征与单纯的几何畸变变得困难。本文提出一种基于近期发表的NAOMI模型(NAOMI model)的原子坐标到分子结构的生成方法。借助这一统一的化学描述框架,本方法即便在输入数据质量不佳的情况下,仍可生成可靠结果。对363条蛋白质数据库(Protein Data Bank, PDB)条目中的分子完成识别,成功率可达98%,这一成果是此前已有方法难以企及的。本方法的鲁棒性通过处理PDB中全部小分子并将其与参考结构比对得到验证。整个数据集的处理耗时不足3分钟,表明本方法适用于大规模应用场景。




