遇见数据集

CimpleG DNAm benchmarking datasets for cell-type classification and deconvolution

收藏
Zenodo2025-09-24 更新2026-05-26 收录
官方服务:

资源简介:

Two large DNAm benchmarking datasets specifically gathered and curated for cell-type classification and deconvolution problems. It includes a leukocytes dataset and a somatic cells dataset in the GenomicRatioSet format from the minfi package. These can be easily loaded into R with the readRDS function: my_data <- readRDS("CimpleG_benchmarking_datasets_2/leukocytes/tidy_leuk_data.rds") Each dataset includes therein sample data like GEO accession numbers, sample name or ID in their original dataset, cell-type label, one-hot encoded data for each cell-type, preferred train/test splits, and others. Alternatively, you can also load the individual .csv files. If you choose this option, I recommend using the function fread from the package data.table. Below I briefly describe these (.csv and .txt) files for the leukocytes dataset, the same logic applies to the somatic cells dataset: tidy_leuk_data_beta-values.csv Methylation Beta values matrix tidy_leuk_data_m-values.csv Methylation M values matrix tidy_leuk_data_probe-annotation.txt Note regarding probe annotation tidy_leuk_data_probe-metadata.csv Probe metadata matrix (chr and location) tidy_leuk_data_sample-metadata.csv Sample metadata matrix (sample ID, cell type labels, one-hot encoded labels, etc.)

提供机构:
Zenodo
创建时间:
2023-06-16
二维码
社区交流群
二维码
科研交流群
商业服务