ccPDB2.0: an updated version of datasets created and compiled from Protein Data Bank
收藏资源简介:
ccPDB 2.0 is an updated and significantly enhanced version of the database of datasets created and compiled from the Protein Data Bank (PDB). This resource provides researchers with high-quality, non-redundant datasets of various protein properties (such as ligand binding, secondary structure, and metal interactions), serving as a gold standard for training and benchmarking machine learning models in structural biology. Web Server: https://webs.iiitd.edu.in/raghava/ccpdb/ Citation Agrawal, P., Patiyal, S., Kumar, R., Kumar, V., Singh, H., Raghav, P. K., & Raghava, G. P. S. (2019). ccPDB 2.0: an updated version of datasets created and compiled from Protein Data Bank. Database, 2019, bay142. https://doi.org/10.1093/database/bay142 About the Research The exponential growth of the Protein Data Bank (PDB) makes it challenging for researchers to manually curate clean, non-redundant datasets for specific structural studies. ccPDB 2.0 automates this process by systematically categorizing protein structures based on their interactions and structural features. Expansion: This version includes 41 different types of datasets, nearly double the amount provided in the initial version. Non-Redundancy: All datasets are curated at multiple sequence identity thresholds (e.g., 25%, 30%, 40%, etc.) to avoid bias during model training.



