idr-sequence-predict datasets: inputs, controls, feature rankings, and consensus region calls
收藏资源简介:
This record includes datasets used by the idr-sequence-predict Snakemake pipeline, designed for proteome-wide identification of intrinsically disordered sequence fragments similar to a small positive set, such as transactivation-domain-like IDRs. This dataset mainly includes the files for the prediction of acidic/hydrophobic short IDRs within TFs across the proteome, based on a small curated dataset. The workflow can expand sequences with sliding windows, compute features using the published idr.mol.feats toolkit, develop control regimes, train multiple ML models, and produce proteome-wide predictions. Predicted fragments are summarised at consensus thresholds, merged into regions, and exported as FASTA files for further analysis. Contents include example inputs, control ID lists, feature importance rankings, selected feature sets, trained models, prediction lists, consensus fragment lists (e.g., 100/95/90/85%), merged region tables, and FASTA sequences of merged regions. The pipeline code is available on GitHub (https://doi.org/10.5281/zenodo.18375255); the feature generator idr.mol.feats is an external dependency (https://github.com/IPritisanac/idr.mol.feats). For reproduction instructions, see the GitHub README for environment setup, configuration, and Snakemake commands.



