遇见数据集

A benchmark dataset of paired wild-type and synthetic genome windows across taxa

收藏
Zenodo2026-03-01 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains paired genomic windows from wild-type (native) genomes and their corresponding synthetic counterparts generated using Evo2, a large-scale evolutionary genome model. For each organism, fixed-length windows were sampled from reference genomes and used as input to Evo2 to generate synthetic sequences that preserve evolutionary and contextual structure while introducing realistic variation. Each pair consists of: an original (wild-type) genomic window extracted from a reference assembly, and a synthetic window produced by Evo2 conditioned on the original sequence. In addition to Evo 2 generated genomes, the dataset now includes a dedicated collection of paired wild-type and synthetic bacteriophage sequences generated using MegaDNA. These bacteriophage pairs consist of fixed-length 50,000 base pair (50 kbp) genomic windows and follow the same pairing structure and directory conventions as the rest of the dataset. The dataset spans multiple taxa, enabling cross-species benchmarking of models that analyze, compare, or learn from genomic sequence data. It is designed to support tasks such as sequence similarity, evolutionary modeling, robustness testing, and evaluation of methods on synthetic but biologically plausible genomes. All sequence pairs are provided as FASTA files, with accompanying CSV files that specify the genomic coordinates, contigs, window sizes, and the mapping between each wild-type and synthetic sequence. This dataset was created in support of the manuscript “Fundamental limitations of genomic language models for realistic sequence generation” and provides the exact data used for the experiments reported in that work.

提供机构:
Zenodo
创建时间:
2026-03-01
二维码
社区交流群
二维码
科研交流群
商业服务