Unified rat-liver microarray expression matrix for 446 DILI-relevant drugs — TG-GATEs + DrugMatrix Affymetrix + DrugMatrix CodeLink + Iconix (GSE8858)
收藏资源简介:
A unified rat-liver gene expression matrix covering 446 DILI-relevant drugs across four toxicogenomic microarray sources, with full pipeline scripts and per-source intermediates for end-to-end reproducibility. Headline numbers 5,859 unique (drug × dose × time) condition rows 24,531 Entrez gene IDs (union of Affymetrix and CodeLink platforms) 446 / 446 drugs covered (100%) Median 4 conditions per drug (range 1–44, mean 13.1) Sources merged SourcePlatformSamplesDrugsOpen TG-GATEs (rat liver in-vivo, single + repeat dose)Affymetrix Rat230_214,143160DrugMatrix-Affymetrix (GEO GSE57815)Affymetrix Rat230_22,218190DrugMatrix-CodeLink (GEO GSE59923)CodeLink Rat UniSet 15,265343Iconix DrugMatrix (GEO GSE8858)CodeLink Rat UniSet 15,312332 Normalization Affymetrix data was processed with RMA + BrainArray ENTREZG v22 (rat2302rnentrezgcdf_22.0.0) — the same recipe used in Eastman & Pande 2019 and our prior reproduction Zenodo deposit 10.5281/zenodo.20204514. Each Affymetrix sample yields 14,132 gene rows (14,075 Entrez IDs + 57 AFFX controls). CodeLink data uses the pre-normalized log2 expression values from the published GEO series matrices (Iconix processing pipeline). GSE8858 has no raw supplementary data on GEO, so uniform re-normalization from raw was not possible for both CodeLink sources. Cross-source merge rules For each (drug, dose, time) condition: (1) Affymetrix wins over CodeLink — if a drug has any Affymetrix data, all of its CodeLink rows are dropped. (2) Within Affymetrix, the same (drug, normalized-dose, normalized-time) measured by multiple Affy sources is averaged. (3) Within CodeLink, the same condition measured in both GSE59923 and GSE8858 is averaged. (4) No drug excluded — all 446 cohort drugs retained. Contents expression_446_per_condition.tsv (607 MB) — main artifact, 5,859 conditions × (9 metadata columns + 24,531 gene columns) expression_446_metadata.tsv (500 KB) — metadata block only expression_446_genes.txt (256 KB) — ordered Entrez gene IDs affy_tggates_single.tsv (693 MB), affy_tggates_repeat.tsv (635 MB), affy_dm_affy.tsv (208 MB) — per-sample log2 RMA matrices (14,132 genes × N samples) per_source/ — per-source .npz + metadata after aggregation but before merge microarray_drug_labels.csv (60 KB) — cohort definition with DILIrank 2.0 labels scripts/ — full pipeline: per-source aggregation Python scripts plus the Neevcloud RMA bundle (install, download, unzip, R RMA script) How to use Each row of expression_446_per_condition.tsv is one ML training example with explicit dose-response and time-course metadata. The platform column lets models account for platform effects; the n_reps_total column gives the underlying replicate count. NaN values appear where a gene is not on that row's platform — CodeLink rows have NaN for ~14K Affymetrix-only genes (and vice versa). Reproducing from raw The Affymetrix pipeline (scripts/neevcloud_affy_rma_pipeline/) is a self-contained bash + R bundle tested on Ubuntu 24.04 with an A6000 GPU instance. From raw CEL ZIPs to the three per-sample log2 RMA TSVs takes about 3 hours including the ~30 GB download. Aggregation and cross-source merge (scripts/01_aggregate_per_source.py + scripts/02_merge_sources.py) run locally in <3 minutes on a laptop. Caveats Affymetrix vs CodeLink expression values are on different scales — audit for platform-effect leakage or apply ComBat/QQ harmonization when pooling. Time strings use each source's native encoding (24 hr, 1 d, 5 days, ...); convert to a single unit (hours) for cross-source matching. Vehicle/control samples are excluded from the per-drug matrix; for log2FC features, subtract matched-control mean per condition using the TG-GATEs AllAttribute control mapping.



