遇见数据集

HADACA3 PDAC MultiOmics Deconvolution Benchmark Datasets

收藏
Zenodo2026-04-30 更新2026-05-26 收录
官方服务:

资源简介:

1. HADACA3 Reference profiles The organizers provide cell population reference profiles for PDAC (pancreatic cancer). Three reference datasets are available: bulk RNAseq of isolated cell populations, bulk methylation profiles of isolated cell populations and single-cell RNAseq. All references are using publicly available data. Each reference type contains 5 cell types: immune cells (immune), fibroblasts (fibro), endothelial cells (endo), and classical (classic) and basal-like (basal) tumor cells. - bulk RNAseq cell populations: immune cells and fibroblasts transcriptomic profiles were retrieved from the GTEx Analysis V10 (GTEx portal, link to paper 1 2), endothelial cell transcriptomic profiles were retrieved from GEO GSE135123 (link to paper). Basal-like and classical tumor cell transcriptomic profiles were retrieved from this paper. Preprocess: this reference data has been normalized using edgeR. - bulk methylation profiles of cell populations: Immune and fibroblasts methylation profiles cells were retrieved from [link to paper](from https://doi.org/10.1186/s12859-021-04381-4). Endothelial cell methylation profiles were retrieved from GSE82234, link to paper. Basal-like and classical cells methylation profiles were retrieved from this paper. Preprocess: this reference dataset consists of aggregated beta value profiles without specific batch effect correction. - single-cell RNAseq: 3 datasets were retrieved. For Peng et al. the 5 cell types are present (link to paper and download). For Baron et al. only endothelial, immune and fibroblast population are present (link to paper and download). For Raghavan et al, basal-like, classical, endothelial and immune cell populations are present (link to paper and download). Preprocess: this reference dataset consists of aggregated beta value profiles without specific batch effect correction. All references are contained in the file ref.h5. This object is a list with the following elements: - ref_bulkRNA: bulk RNAseq pure cell populations - ref_met: bulk methylation profiles of pure cell populations - ref_scRNA: single-cell RNAseq reference datasets from 3 differents studies; Peng, Baron and Raghavan, each including the data (counts) and associated metadata, i.e. the sample and the cell type. Example to load and inspect the reference data: 2. HADACA3 Benchmark Datasets This repository provides the datasets used in the HADACA3 benchmark, designed to evaluate multi-omic deconvolution methods across diverse biological and simulated settings. 2.1 Overview The benchmark includes 9 datasets covering three scenarios:- In vivo (real tumor samples)- In vitro (controlled mixtures)- In silico (simulated data) All datasets include matched RNA-seq and DNA methylation (DNAm) profiles, restricted to a common feature space (~20k RNA features, ~23k DNAm features). 2.2 Datasets Name Type File Name Samples Cell types Description VIVO In vivo invivo 47 2 PDAC tumor samples with histology-derived proportions VITR In vitro invitro 30 5 Controlled mixtures with known proportions SBN5 In silico insilicopseudobulk 60 5 Pseudo-bulk from simulated single-cell data SDN5 In silico insilicodirichletNoDep 60 5 Heteroscedastic noise SDN4 In silico insilicodirichletNoDep4CTsource 60 4 Missing cell types SDN6 In silico insilicodirichletNoDep4CTsource 60 6 Extra cell types SDE5 In silico insilicodirichletEMFA 60 5 Structured noise (EM) SDEL In silico insilicodirichletEMFAImmuneLowProp 60 5 Rare cell types SDC5 In silico insilicodirichletCopule 60 5 Correlated features (copula-based) 2.3 Data generation Simulated datasets are generated using a linear mixture model, combining cell-type reference profiles with sampled proportions and modality-specific noise. RNA and DNAm share the same underlying proportions. The SBN5 dataset*uses pseudo-bulk aggregation to better mimic realistic experimental variability. Reference profiles used for deconvolution are independent from those used to generate simulated mixtures, ensuring fair evaluation. 2.4 Usage These datasets are intended for:- Benchmarking deconvolution methods- Evaluating multi-omic integration strategies- Reproducible research in computational biology Dataset metadata is available in Croissant format (croissant_pdac_deconvolution.json), compatible with mlcroissant and the Hugging Face datasets hub. For more details, please refer to the associated publication. Datasets were split into two replicates. The first replicate serves to test models, while the second was included to prevent overfitting on the first dataset HDF5 format is convenient for maintaining read and write performance while preserving compatibility with sparse matrices, dataframes, and lists, and remaining friendly to both R and Python. 3. HADACA3 Benchmark and Competition Results In addition to the datasets, we provide the results of the modular benchmark used to generate the figures presented in the associated publication. These results correspond to the evaluation of hundreds of thousands of pipeline combinations across all benchmark datasets and are provided as compressed CSV files. 3.1 Files The HADACA3 benchmark results are split into six files, according to:- the integration strategy: early integration (EI) or late integration (LI)- the dataset groups / simulation scenarios Early integration (EI)- `results_ei_EMFA_Copule_Immuneprop1.csv.gz` - `results_ei_invitro_invivo_insilicopsuedobulk.csv.gz` - `results_ei_NODEP6_NODEP4_NODEP1.csv.gz` Late integration (LI)- `results_li_EMFA_Copule_Immuneprop1.csv.gz` - `results_li_invitro_invivo_insilicopsuedobulk.csv.gz` - `results_li_NODEP6_NODEP4_NODEP1.csv.gz` The HADACA3 competition result corresponding to JOKER method results is named HADACA3_TeamJ_Results.zip 3.2 Content Each file contains:- the pipeline configuration (preprocessing, feature selection, deconvolution, integration)- the dataset identifier- the evaluation metrics, including the aggregated benchmark score - Files are provided in compressed `.csv.gz` format for efficiency.- Results are split to reduce file size and improve usability.- These files allow full reproducibility of the figures and analyses presented in the paper.

提供机构:
Zenodo
创建时间:
2026-04-30
二维码
社区交流群
二维码
科研交流群
商业服务