Inflammatory bowel diseases (IBD) gene expression data for the manuscript entitled "TENTACLES: a consensus Machine Learning tool for robust biomarker discovery in heterogeneous data"
收藏资源简介:
RNA sequencing reads related to IBD intestinal biopsies were obtained from the Sequence Read Archive (SRA). All samples were uniformly processed using a standardized custom pipeline to ensure consistency across datasets. Quality control was performed with fastp (v1.0.1), followed by alignment to the human reference genome GRCh38.p14 (NCBI RefSeq assembly GCF_000001405.40) using STAR (v2.7.10) with default parameters. Read strandedness was assessed with RSeQC (v5.0.1), and transcript quantification was carried out using Rsubread (v2.22.1). Publicly available datasets included in this repository: PRJNA248469, PRJNA565216, PRJNA702434, PRJNA728646, PRJNA985602, PRJNA961704. ibd_datasets.rar – contains gene expression count data and clinical metadata from all included studies. PRJNA248469 – Dataset 1 (d1) (Ensemble-based Signature Discovery Cohort) PRJNA565216 and PRJNA985602 – Dataset 2 (d2) (Cross-cohort Signature Evaluation Cohort) PRJNA702434 – Dataset 3 (d3) (Minimal Signature Identification Cohort) PRJNA728646 and PRJNA961704 - Dataset 4 (d4) (Independent Unsupervised Validation Cohort) ibd_datasets.Rdata - contains all tables in a single .Rdata object. TENTACLES_manuscript_vignette.html - contains a demonstration vignette that shows how to work with TENTACLES. ibd_datasets.Rdata is used for this demonstration. TENTACLES_manuscript_vignette.pdf - contains the same demonstration vignette but in PDF format.*Current version: Revised in response to peer-review feedback.



