遇见数据集

idr-sequence-predict datasets: inputs, controls, feature rankings, and consensus region calls

收藏
Zenodo2026-03-18 更新2026-05-26 收录
官方服务:

资源简介:

This record includes datasets used by the idr-sequence-predict Snakemake pipeline, designed for proteome-wide identification of intrinsically disordered sequence fragments similar to a small positive set, such as transactivation-domain-like IDRs. This dataset mainly includes the files for the prediction of acidic/hydrophobic short IDRs within TFs across the proteome, based on a small curated dataset. The workflow can expand sequences with sliding windows, compute features using the published idr.mol.feats toolkit, develop control regimes, train multiple ML models, and produce proteome-wide predictions. Predicted fragments are summarised at consensus thresholds, merged into regions, and exported as FASTA files for further analysis. Contents include example inputs, control ID lists, feature importance rankings, selected feature sets, trained models, prediction lists, consensus fragment lists (e.g., 100/95/90/85%), merged region tables, and FASTA sequences of merged regions. The pipeline code is available on GitHub (https://doi.org/10.5281/zenodo.18375255); the feature generator idr.mol.feats is an external dependency (https://github.com/IPritisanac/idr.mol.feats). For reproduction instructions, see the GitHub README for environment setup, configuration, and Snakemake commands.

提供机构:
Zenodo
创建时间:
2026-03-18
二维码
社区交流群
二维码
科研交流群
商业服务