遇见数据集

ModCon: A unified framework for quantifying epitranscriptomic conservation and prioritizing functional RNA modification sites

收藏
Zenodo2026-07-10 更新2026-08-01 收录
官方服务:

资源简介:

ModCon is a unified framework for quantifying epitranscriptomic conservation at single-base resolution. Unlike conventional conservation methods that measure nucleotide sequence conservation, ModCon evaluates whether an RNA modification event itself is evolutionarily retained across species. Built from Oxford Nanopore direct RNA sequencing (ONT) datasets across human, mouse, and pig, ModCon integrates four complementary evidence components to generate a unified conservation score for every modified residue. These complementary signals include orthologous modification concordance, within-species recurrence, local sequence-context conservation, and a fine-tuned DNABERT-2 language-model-derived conservation score. The resulting ModCon score (0–1) provides a quantitative measure of evolutionary conservation at RNA modification sites, enabling researchers to prioritize conserved and potentially functional RNA modifications for downstream biological analyses. Dataset for ModCon 1. Reference genome and Chains: hg38ToHg19.over.chain —— Human genome liftover chain file (GRCh38 to GRCh37). hg38ToMm39.over.chain —— Cross-species liftover chain file (Human GRCh38 to Mouse GRCm39). hg38ToSusScr11.over.chain —— Cross-species liftover chain file (Human GRCh38 to Pig Sscrofa11.1). hg38_chr1.fa —— Human reference genome sequence (UCSC version, GRCh38). mm39_chr1.fa —— Mouse reference genome sequence (UCSC version, GRCm39). susScr11._chr1.fa —— Pig reference genome sequence (UCSC version, SGSC Sscrofa11.1). Note: Due to reference genome is too large, here I use chr1.fa as example, please replace it to whole reference genome. 2. Species Data: Includes 1-based coordinate modification sites derived from the Oxford Nanopore Technologies (ONT) pipeline. ontdata —— All human ONT-derived RNA modification sites. mouse_raw —— All mouse ONT-derived RNA modification sites. pig_raw ——All pig ONT-derived RNA modification sites. 3. NGS data: Publicly available high-throughput Next-Generation Sequencing (NGS) modification datastes. Human(m6A,m1A,m5C,ac4C,m7G,m6Am, Am,Cm,Gm,Um,m5U,psi) Mouse (m6A,m1A,m5C) Rat (m6A) Zebrafish (m6A) SomaticSNP —— Human cancer-associated somatic mutations (e.g., from TCGA/gnomAD). GermlineSNP —— Human cancer-associated germline mutations (e.g., from TCGA/gnomAD). 4. Training Data: Data extracted from negative and positive data pools used for model training. Model_A_data —— Training data for ModCon-A. Model_C_data —— Training data for ModCon-C. Model_G_data —— Training data for ModCon-G. Model_U_data —— Training data for ModCon-U. 5. Sub-base Models:Pretrained base-specific models of ModCon to predict modification conservation levels. ModCon-A —— Pretrained model for the A base. ModCon-C —— Pretrained model for the C base. ModCon-G —— Pretrained model for the G base. ModCon-U —— Pretrained model for the U base. 6.Functional validation data: hg38_transcript —— 20024 human primary transcripts. RBP_site_hg38 —— RNA-Binding Protein (RBP) binding peaks. SplingSite_hg38_500bp —— Splice sites (5' and 3' SS) expanded with a ±500bp flanking genomic window. Code: 1. Oxford Nanopore data processing: The raw data processing pipline includes base-calling, alignment, and moficiation detection. Raw data processing pipline.sh ——The raw ONT data processing pipline. 2.Model training and SHAP anlysis: train.py —— Integrated DNABERT2 big language model+MLP fine-tuning code. SHAP_analysis.py —— The SHAP analysis code to illustrate feature importance. 3. Assessment of each evidence component: To test whether each evidence components carry independently informative conservation signals. orthologous_cocordance.Rmd sequence-context_conservation.Rmd Within_species_recurrence.Rmd 4.Integration of ModCon evidence:We integrated these signals into a quantitative ModCon score and tested whether their integration improves site-level prioritization. pairwise.Rmd —— The spearman correlation of each component. Performance of each evidence.py —— The AUROC,AUPRC and top10 enrichment value of each component. Bootstrap comparison.py —— The bootstrap with replacement of each component. 5.NGS validation code: To test whether ModCon generalizes beyond the ONT data on which it was built, we collected independent base-resolution modification datasets generated by NGS-based methods. To test whether ModCon could identify cross-species conserved modification sites (CSCM-NGS) and show genetic constraint within human populations. m1A_CSCM.Rmd —— The m1A CSCM test. m5C_CSCM.Rmd —— The m5C CSCM test. m6A_CSCM.Rmd —— The m6A CSCM test. SNP_variant_density_and_deleterious_level.Rmd —— Including the Somatic,Germline variants ratio and related deleterious level. 6.Functional analysis: Downstream Biological Insights RNA_Binding_protein.Rmd —— The effect of the level of modification conservation on RNA binding proteins. Splicing_sites.Rmd —— Evaluates potential regulatory roles of different conserved modification levels near splice junctions. Clustering.Rmd —— clustering effects across different conservation levels Gene_ontology.Rmd —— Performs Gene Ontology (GO) and KEGG pathway enrichment analysis on different conservation level groups to uncover their potential biological roles. Note:For comprehensive usage, instructions please visit our GitHub repository: https://github.com/Gang1998c/Modcon

提供机构:
Zenodo
创建时间:
2026-07-10
二维码
社区交流群
二维码
科研交流群
商业服务