Transcriptional co-expression and sequence-model scoring of the Prochlorococcus MED4 CrY2H-seq interactome: a 227-run RNA-seq analysis with functional, WGCNA, and model-based cross-checks
收藏资源简介:
Analysis scripts, intermediate data, final figures, and a PDF writeup for two orthogonal analyses of the Prochlorococcus marinus MED4 protein-protein interaction network determined by CrY2H-seq, supporting a manuscript in preparation on the MED4 interactome. (1) Transcriptional co-expression: whether the experimental interactome, and its high-betweenness bottleneck nodes (including short open reading frames absent from the 2003 reference annotation), is co-expressed above a degree-preserving null across a uniformly re-quantified compendium of 227 public RNA-seq runs (7 studies with ≥6 samples; 216 samples), using Spearman correlation with CLR background normalization, plus comprehensive per-node/community, GO/InterPro, and WGCNA cross-checks and de-novo InterProScan of the hub proteins. (2) Sequence-model scoring: three trained sequence classifiers (Siamese ESM-1b, an ESM cross-encoder, and a character-level GPT) score the 17,116 low-confidence candidate pairs; a three-model consensus at P(interact)≥0.95 endorses 552 pairs (3.2%) as a model-based shortlist. Because the models were trained on the high-confidence positives, this is model endorsement/prediction, not independent validation. Contents: analysis scripts; intermediate data (network, ID crosswalk, short-ORF-augmented annotation, count matrix and sample metadata, co-expression result tables, InterProScan results, per-model score tables); final figures; and a typeset PDF writeup with LaTeX source. Raw sequencing reads and model checkpoints are NOT redeposited here (reads remain in SRA/ENA by accession; checkpoints are on Hugging Face). Code MIT; derived data CC-BY-4.0.



