Inflammatory bowel diseases (IBD) gene expression data for the manuscript entitled "Supervised Ensemble Learning Identifies Minimal Consensus Gene Signatures for Crohn's Disease Classification"
收藏资源简介:
RNA sequencing reads related to IBD intestinal biopsies were obtained from the Sequence Read Archive (SRA). All samples were uniformly processed using a standardized custom pipeline to ensure consistency across datasets. Quality control was performed with fastp (v1.0.1), followed by alignment to the human reference genome GRCh38.p14 (NCBI RefSeq assembly GCF_000001405.40) using STAR (v2.7.10) with default parameters. Read strandedness was assessed with RSeQC (v5.0.1), and transcript quantification was carried out using Rsubread (v2.22.1). Publicly available datasets included in this repository: PRJNA248469, PRJNA565216, PRJNA702434, PRJNA728646, PRJNA985602. ibd_counts.csv – contains gene expression count data from all included studies. To avoid non-interpretable results and reduce computation time, genes not annotated in GO or KEGG were filtered out from this table. ibd_clin.csv – contains clinical metadata for all samples across all studies. ibd_dataset.Rdata - contains both tables in a single .Rdata object. TENTACLES_Workflow.html - contains a demonstration vignette that shows how to work with TENTACLES. ibd_dataset.Rdata is used for this demonstration.



