pydangle top100: Per-residue backbone geometry for 106 quality-filtered protein chains from the original Richardson Lab reference dataset
收藏资源简介:
Per-residue backbone geometry (phi, psi, omega, tau), Ramachandran classifications at four granularities (6/5/4/3-class), DSSP secondary structure, peptide bond type, and chirality for 17,434 quality-filtered protein residues from 106 chains in 98 high-resolution protein structures. Computed using pydangle-biopython v0.5.1 on "ersatz" full-structure PDB files with NQH flip corrections from Reduce 4.16. Source data from the Richardson Lab Top100 reference dataset, the original quality-filtered set used for foundational NQH flip correction work (Word et al., 1999, doi:10.1006/jmbi.1998.2401). Residue-level filtering by mainchain B-factor <= 40, matching the original Top100 methodology (the larger Top500/Top8000 datasets use the stricter B <= 30 threshold). See README.md for full methodology.
本数据集包含98个高分辨率蛋白质结构的106条链中筛选得到的17434个经质量筛选的蛋白质残基,其标注项涵盖:每残基主链几何参数(phi、psi、omega、tau)、四种粒度(6/5/4/3类)的拉马钱德兰(Ramachandran)分类、DSSP二级结构、肽键类型与残基手性。所有特征均通过pydangle-biopython v0.5.1,基于搭载Reduce 4.16版本NQH翻转校正的"ersatz"(模拟完整结构)完整结构PDB文件计算得到。数据集原始数据源自理查森实验室(Richardson Lab)的Top100参考数据集,该数据集是支撑奠基性NQH翻转校正研究的原始质量筛选集(Word 等,1999,doi:10.1006/jmbi.1998.2401)。残基级筛选标准为主链B因子≤40,与原始Top100数据集的处理流程一致(规模更大的Top500/Top8000数据集采用了更为严格的B因子≤30阈值)。完整实验方法详见README.md文件。



