TriZOD — Re-referenced BMRB chemical-shift dataset for protein disorder
收藏资源简介:
Per-residue protein-disorder labels derived from BMRB NMR backbone chemical shifts. Observed shifts are re-referenced (LACS + POTENCI/AIC) and scored against POTENCI random-coil predictions to give CheZOD-style per-residue Z-scores and G-scores (0-1, higher = more disordered). The dataset is distributed as a SINGLE Parquet file (trizod_dataset.parquet): one row per scored protein chain, with per-residue Z/G-score, shift-count (k) and boolean mask arrays stored as list columns aligned 1:1 to the sequence, plus NMR sample conditions, re-referencing offsets and citation provenance. Four nested stringency tiers (unfiltered/tolerant/moderate/strict) and redundancy-reduced training representatives are selectable via the split and train_tier columns; two held-out, leakage-free test sets are included (CheZOD117, 115 chains; TriZOD, 344 chains). Produced by the finalized TriZOD pipeline (2026-06 build).



