遇见数据集

Dataset for "M6AFormer Prioritizes Unannotated Functional m6A Candidate Sites in the Human m6A Epitranscriptome"

收藏
Zenodo2026-07-10 更新2026-08-01 收录
官方服务:

资源简介:

M6AFormer transcriptome-wide high-confidence m6A predictions and functional enrichment tables This record provides the transcriptome-wide m6A site predictions and the associated functional and disease enrichment tables generated by M6AFormer, a sequence-based model for prioritizing candidate N6-methyladenosine (m6A) sites across the human transcriptome. Overview M6AFormer was applied to the hg38 reference to score every candidate adenosine in annotated exonic transcript regions. Because a genome-wide scan evaluates every adenosine rather than only motif-restricted sites, scanning used the all-adenosine 801-nt model (ALL801), whose negative training distribution matches the full genomic adenosine population. Sites passing a stringent high-confidence probability cutoff of prob >= 0.99 are released here. Each site is classified as annotated when it overlaps the union of curated m6A resources (RMBase and m6A-Atlas) and as unannotated otherwise. The unannotated high-confidence set represents candidate m6A sites that are not yet captured by current annotation databases but resemble known m6A sites in sequence and regulatory context. Contents The record contains two data groups and a README data dictionary: 1. High-confidence predicted m6A sites (prob > 0.99)** A single table of 312,625 unique genomic loci (232,128 annotated, 80,497 unannotated), with genomic position, strand, predicted probability, annotation status, DRACH motif status, sequence motifs, the 101-nt sequence context, and full transcript-level annotation (gene, transcript, region, relative position, and distance to start and stop codons). 2. Functional and disease enrichment tables** Over-representation results (clusterProfiler) against Gene Ontology Biological Process (GO BP), KEGG, Disease Ontology (DO), and DisGeNET (DGN), provided separately for the annotated and the unannotated (prob > 0.99) gene sets. The annotated background uses the complete set of database-verified m6A sites, while the unannotated set uses the high-confidence predictions defined above. File manifest - M6AFormer_high_confidence_m6A_sites_prob_ge_0.99.tsv.gz - GO_BP_annotated.csv - GO_BP_unannotated_prob_gt_0.99.csv` - KEGG_annotated.csv` - KEGG_unannotated_prob_gt_0.99.csv` - DO_annotated.csv` - DO_unannotated_prob_gt_0.99.csv` - DGN_annotated.csv` - DGN_unannotated_prob_gt_0.99.csv` Notes Genomic coordinates are 1-based on the hg38 assembly. Predicted probabilities are model outputs and are not direct measurements of methylation, so high-confidence sites still require experimental validation. Full methods are described in the associated publication.

提供机构:
Zenodo
创建时间:
2026-07-10
二维码
社区交流群
二维码
科研交流群
商业服务