An alphabet of allostery
收藏资源简介:
An Alphabet of Allostery An Alphabet of Allostery: a transferable transfer entropy contact vocabulary and pooled pattern database Description An Alphabet of Allostery is a transferable, zero-fit transfer entropy (TE) contact vocabulary that predicts allosteric-site character and conformational rewiring from a single protein structure. Gaussian Network Model (GNM) dynamics are computed over a non-redundant survey of the Protein Data Bank; each structure is decomposed into local contact words anchored by spatial cliques, and the per-word net TE statistics are pooled into an alphabet. Any new structure can then be projected onto that alphabet to yield a per-residue allosteric track with four channels: net TE, source, sink, and switch. Sink residues are transfer entropy receivers and mark annotated allosteric sites, validated against the Allosteric Database (ASD). Source residues are drivers and predict where a protein rewires between conformational states, validated on nine two-state apo-to-holo pairs. Switch residues track conformational hinges and seed a candidate design vocabulary. Honest scope: the effects are real and correctly signed but modest (ASD sink ROC-AUC around 0.54; per-protein rewiring Spearman rho around 0.1). The contribution is a universal, parameter-free prior built from the physics of information flow, not a high-accuracy site predictor. This deposition contains the analysis code (an importable scoring method plus command-line tools for scoring any PDB chain and reproducing the paper results), the ASD validation and figure-generation scripts, the full generation pipeline that builds the alphabet, and the published figures, per-protein tracks, scored structures, and metric tables. The pooled pattern database itself (a ~3.9 GB Parquet file) is included as a separate archive component; the code auto-discovers it under data/ when unzipped. Built from ~68,000 non-redundant PDB chains (PISCES), the pooled alphabet comprises 131,611,766 unique contact words from 212,860,934 clique observations. Version v1.0.0 (8.0 Å clique-cutoff build) Contents component notes src/, scripts/ , analysis/, generation/, results/ TE method, CLI tools, analysis tools, generation of the data, results(all zipped) pooled_pattern_database.parquet 3,939,723,037 bytes; SHA-256 3a5391ad57cadfc08d327953e1b5c29272dfcd29ff0c938bcd341d340dcafb9e; schema Pattern, Mean_TE_Score, Std_TE_Score, Count; 131,611,766 rows README.md, REPRODUCE.md full usage and end-to-end reproduction recipe (travel inside the archive)



