遇见数据集

Datasets for DuET: A Unified Deep Learning Framework for Predicting mRNA Translation Efficiency Across Human Cell Types

收藏
Zenodo2026-07-14 更新2026-08-02 收录
官方服务:

资源简介:

This deposit contains the training, benchmark, and reference datasets for DuET, a deep learning framework that predicts mRNA translation efficiency (TE) from 5' UTR and CDS sequences jointly across 76 human cell types. Translation efficiency values are derived from paired ribosome-profiling and RNA-seq data. All sequence tables are tab-separated (`.tsv`) with transcript identifiers (`txID`), 5' UTR / CDS sequences, and TE labels. Total uncompressed size is approximately 3.4 GB. Extract all tar.gz archives and place them under the datasets/ directory in the duet source directory. Contents: datasets/├── multitask_TE.tsv├── sequence_features.tsv├── celltype_te/│ ├── all-celltype_TE.tsv│ ├── a549_TE.tsv│ ├── hek293t_TE.tsv│ └── ... (76 files total)├── mrl_benchmark/│ ├── U_1.tsv, U_2.tsv, pU_1.tsv, pU_2.tsv, m1pU_1.tsv, m1pU_2.tsv, mC-U_1.tsv, mC-U_2.tsv│ ├── train_*.tsv│ ├── *_sequence_features.tsv│ └── cds.fasta└── transcriptome/ ├── duet_transcriptome.selected.fa ├── duet_transcriptome.selected.gtf ├── duet_transcriptome.selected.tsv └── duet_transcriptome_valid_window.saf multitask_TE.tsv — Multi-task training table: one row per gene with 5'UTR and CDS sequences and 76 cell-type TE columns (TE_*). sequence_features.tsv — Precomputed sequence features (MFE, GC content, Kozak similarity, start-context statistics, etc.) used by the feature-basedmodels. celltype_te/ — Single-task per-cell-type TE tables (~3.1 GB), one file per cell type. Includes all-celltype_TE.tsv (pan-cell-type averaged TE, used to train the default checkpoint) and 75 individual human cell types / tissues (cell lines, primary cells, iPSC/ESC-derived cells, and tissues). mrl_benchmark/ — Mean ribosome load (MRL) benchmark from a polysome-profiling MPRA of modified 5'UTRs. Test sets are {U,pU,m1pU,mC-U}_{1,2}.tsv and matching training splits are train_*.tsv, with *_sequence_features.tsv giving precomputed features per split and cds.fasta the shared CDS context. transcriptome/ — DuET transcriptome definition: selected transcript sequences (.fa), their annotation (.gtf), a transcript metadata table (.tsv), and valid window intervals for featureCounts (.saf). Nucleotide chemistries in mrl_benchmark/: U (uridine), pU (pseudouridine), m1pU (N1-methylpseudouridine). The U, pU, and m1pU sets use an eGFP reporterCDS, while mC-U uses uridine with an mCherry reporter CDS.

提供机构:
Zenodo
创建时间:
2026-07-14
二维码
社区交流群
二维码
科研交流群
商业服务