TRACE: Deciphering Translation Grammar to Predict Translation Dynamics from RNA Sequences
收藏资源简介:
TRACE (Translatome Representation Across Cell Environment) is a Transformer-based model that decodes full-length transcriptomes into translatomes — predicting per-position ribosome occupancy profiles directly from RNA sequence and cellular context. TRACE integrates multi-omics data — transcript sequence, gene expression, and species identity — through adaptive layer normalization (AdaLN-Zero) to resolve translation regulation across cell types and species. Pre-processed training/validation/test H5 datasets and expression dictionariesare publicly available: > The archive contains `.train.h5` / `.valid.h5` / `.test.h5` files for each> species (human, macaque, mouse) as well as per-species expression dictionaries> (`{species}_expression_dict.pt`). ### H5 File Layout ```<dataset>.h5├── .attrs["n_samples"] int total number of transcripts├── .attrs["cell_type_counts"] JSON str cell-type distribution├── /cell_exprs/│ └── <cell_type> (d_expr,) Z-scored expression vector per cell type├── /sequences/│ └── <tid> (L, d_seq) continuous sequence features (e.g., one-hot codon)└── /samples/<uuid>/ ├── .attrs["species"] str human | macaque | mouse ├── .attrs["cell_type"] str e.g., "heart", "liver", "HepG2" ├── .attrs["cds_start_pos"] int16 CDS start (1-based), -1 if unknown ├── .attrs["cds_end_pos"] int16 CDS end (1-based), -1 if unknown ├── .attrs["te_scale"] float32 translation efficiency (Z-scored), None if missing ├── .attrs["rpf_depth"] float32 ribosome profiling depth ├── .attrs["rpf_coverage"] float32 ribosome profiling coverage ├── .attrs["motif_occ"] list[int] upstream motif occurrences └── count_emb (L, d_count) per-position RPF density (target)``` ### Loading Data The `TranslationDataset` class provides lazy (recommended) or eager loadingfrom `.h5` files: ```pythonfrom data.translation_dataset import TranslationDataset # Lazy loading — minimal RAM, reads on-demand per __getitem__ds = TranslationDataset.from_h5("human.train.h5", lazy=True) print(f"Samples: {ds.n_samples}")print(f"Cell types: {ds.cell_type_counts}")print(f"Cell expr dict keys: {list(ds.cell_expr_dict.keys())}") # Access a single sample (returns one transcript)tid, species, cell_type, expr_vector, meta, seq_emb, count_emb = ds[0] print(f"species: {species}") # e.g., "human"print(f"cell_type: {cell_type}") # e.g., "heart"print(f"seq_emb shape: {seq_emb.shape}") # (L, d_seq)print(f"count_emb shape: {count_emb.shape}") # (L, d_count)print(f"cds: {meta['cds_start_pos']}–{meta['cds_end_pos']}")print(f"te_scale: {meta.get('te_scale')}") # Expression vector for a specific cell typeexpr_vec = ds.cell_expr_dict["heart"] # shape (d_expr,)```



