遇见数据集

Comparing machine learning methods predicting transcriptome from epigenome with applications to association studies

收藏
Zenodo2026-05-20 更新2026-05-26 收录
官方服务:

资源简介:

Understanding how epigenome variation contributes to gene expression in disease and development is a fundamental challenge. Regulatory regions show cell type-specific epigenome activity and differ in their location, size, and distance to their target genes, complicating discovery and analysis. Recent machine learning models have been proposed to address these problems by learning functions for the prediction of gene expression from epigenomic data. Here, we use the large IHEC EpiATLAS dataset to benchmark state-of-the-art linear and non-linear approaches. Each approach is optimized for over 28,000 human genes, providing a comprehensive regulatory catalog of gene models. In-depth comparison reveals that gene characteristics and the epigenomic complexity of the locus influence the difficulty of predicting the epigenome-to-transcriptome association. The model performance is further evaluated using CRISPRi and eQTL validation data. Based on these models, we conduct histone-acetylation association studies in a systematic way to investigate how epigenomic variation impacts gene expression. The model-based analysis revealed genes and regulatory regions linked to B-cell leukemia in patient data with known disease-related functions.Our work provides a foundation for applications that link epigenome variation to gene expression in human cells, by benchmarking methods on a per-gene basis, illustrating their use in a disease context and making trained models available to the community. Our codes for the analysis are here and below is files description: AML_CLL_disgenet_datasets.txt : CLL related genes from DisGenet dataset CNN_best_models.tar.gz : Binned-CNN trained models CRE_RF_models.tar.xz : CRE-RF trained models disease_samples_12.txt : 12 CLL diases samples healthy_samples_10.txt : 10 CLL healthy samples genes_expected_count_DESeq2_H3K27acFormatted.tsv : RNA-seq counts from IHEC after matching RNA-seq and H3K27ac ChIP-seq experiments per EpiRR ID eQTLs_MinimumOverlap.txt : Minimum number of interactions with eQTL-gene overlaps across methods IHEC_H3K27ac_RNA_uuid_mapping.tsv : Mapping of the RNA-seq uuid (experiment IDs) to the H3K27ac ChIP-seq uuids IHEC_metadata_Grade1ColonAdenocarcinoma.tsv : The metadata for Grade1ColonAdenocarcinoma disease JointValidatedInteractions_hg38_500kb.txt : Table with CRISPRi-validated enhancer-gene interactions limited to a maximum distance of 500 kb partition0_test.csv : Original test partition (name of the samples) partition0_train.csv : Original train partition (name of the samples) Provided_Input.zip : Folder with the required files to generate predictions on custom data with the code in the EpiExpress repository Review0725_CTSGenes_SubsetCTS_211025.txt : List of genes that were used for cell type-specific partiotion training/validation test_partition_distinct_celltypes.csv : Test samples for the cell type-specific partition train_partition_distinct_celltypes.csv : Train samples for the cell type-specific partition coloncancer_samples.txt: list of the samples were used for the HAWAS analysis for the colon cancer CRE_input.tar.gz: Input files for traing CRE methods (Due to its size, we are unable to provide the input for the Binned models. They can be reproduced by getting the bigwig files of the train and test samples from the IHEC data portal, and then use the EpiExpress repository to generate the matrices.)

提供机构:
Zenodo
创建时间:
2026-05-19
二维码
社区交流群
二维码
科研交流群
商业服务