遇见数据集

ccPDB2.0: an updated version of datasets created and compiled from Protein Data Bank

收藏
Zenodo2026-05-11 更新2026-05-26 收录
官方服务:

资源简介:

ccPDB 2.0 is an updated and significantly enhanced version of the database of datasets created and compiled from the Protein Data Bank (PDB). This resource provides researchers with high-quality, non-redundant datasets of various protein properties (such as ligand binding, secondary structure, and metal interactions), serving as a gold standard for training and benchmarking machine learning models in structural biology. Web Server: https://webs.iiitd.edu.in/raghava/ccpdb/ Citation Agrawal, P., Patiyal, S., Kumar, R., Kumar, V., Singh, H., Raghav, P. K., & Raghava, G. P. S. (2019). ccPDB 2.0: an updated version of datasets created and compiled from Protein Data Bank. Database, 2019, bay142. https://doi.org/10.1093/database/bay142 About the Research The exponential growth of the Protein Data Bank (PDB) makes it challenging for researchers to manually curate clean, non-redundant datasets for specific structural studies. ccPDB 2.0 automates this process by systematically categorizing protein structures based on their interactions and structural features. Expansion: This version includes 41 different types of datasets, nearly double the amount provided in the initial version. Non-Redundancy: All datasets are curated at multiple sequence identity thresholds (e.g., 25%, 30%, 40%, etc.) to avoid bias during model training.

提供机构:
Zenodo
创建时间:
2026-05-11
二维码
社区交流群
二维码
科研交流群
商业服务