A Structural Geometry Dataset for Proteins: Distance Maps and Dihedral Angles for Computational Applications
收藏资源简介:
This dataset provides a comprehensive collection of geometric features for 6,464 single-chain proteins, including precomputed distance maps and dihedral angles. The dataset was curated from the SCOPe 2.08 database, focusing on proteins with chain lengths between 25 to 700 residues, ensuring diversity while avoiding redundancy. Key Features: Distance Maps: The dataset includes pairwise Cβ-Cβ (or Cα for Glycine) distance matrices, which are essential for understanding protein tertiary structure, interactions, and folding. The maps are generated using the mean squared error (MSE) metric, providing a 2D representation of inter-residue distances. Dihedral Angles: Backbone torsion angles (ϕ, ψ, and ω) are calculated based on atomic positions of Cα, N, and C atoms. These angles are vital for describing protein conformation, secondary structure (e.g., alpha-helices and beta-sheets), and stability. Handling of Missing Residues: The dataset addresses missing residues from PDB files through interpolation for small gaps (1-2 residues) and masking for larger gaps, ensuring accurate structural data. File Format: The geometric data is stored in compressed .npz files, which include the distance matrix, dihedral angle vectors, sequence indices, resolution classification, and a mask for missing residues. Applications: This dataset is intended for a range of structural bioinformatics applications, including: Protein Structure Prediction: Distance maps and dihedral angles can serve as inputs for machine learning models that predict protein structures from sequences. Molecular Dynamics Simulations: The dataset provides geometric constraints that can enhance molecular simulations and improve conformational sampling. Structural Analysis: The dataset facilitates comparative studies of protein families, structural motifs, and evolutionary relationships. Drug Discovery: The geometric features can aid in ligand binding and protein-ligand docking studies, which are crucial for structure-based drug design. By making this dataset publicly available, we aim to support advancements in protein structure prediction, molecular simulations, and AI-driven bioinformatics research.



