遇见数据集

A benchmark dataset of paired wild-type and synthetic genome windows across taxa

收藏
Zenodo2026-03-01 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains paired genomic windows from wild-type (native) genomes and their corresponding synthetic counterparts generated using Evo2, a large-scale evolutionary genome model. For each organism, fixed-length windows were sampled from reference genomes and used as input to Evo2 to generate synthetic sequences that preserve evolutionary and contextual structure while introducing realistic variation. Each pair consists of: an original (wild-type) genomic window extracted from a reference assembly, and a synthetic window produced by Evo2 conditioned on the original sequence. In addition to Evo 2 generated genomes, the dataset now includes a dedicated collection of paired wild-type and synthetic bacteriophage sequences generated using MegaDNA. These bacteriophage pairs consist of fixed-length 50,000 base pair (50 kbp) genomic windows and follow the same pairing structure and directory conventions as the rest of the dataset. The dataset spans multiple taxa, enabling cross-species benchmarking of models that analyze, compare, or learn from genomic sequence data. It is designed to support tasks such as sequence similarity, evolutionary modeling, robustness testing, and evaluation of methods on synthetic but biologically plausible genomes. All sequence pairs are provided as FASTA files, with accompanying CSV files that specify the genomic coordinates, contigs, window sizes, and the mapping between each wild-type and synthetic sequence. This dataset was created in support of the manuscript “Fundamental limitations of genomic language models for realistic sequence generation” and provides the exact data used for the experiments reported in that work.

本数据集包含野生型(wild-type)基因组的成对基因组窗口,以及借助大规模进化基因组模型Evo2生成的对应合成序列。针对每一种生物体,研究人员从参考基因组中采样固定长度的窗口序列作为Evo2的输入,以此生成既保留进化与上下文结构、又引入真实变异的合成序列。 每一对样本包含: - 从参考基因组组装结果中提取的原始野生型基因组窗口; - 以原始序列为条件,由Evo2生成的合成窗口。 除Evo2生成的基因组外,本数据集还新增了一批专用的成对野生型与合成噬菌体序列,这些序列由MegaDNA生成。该类噬菌体成对样本采用固定长度为50000碱基对(50 kbp)的基因组窗口,且与数据集其余部分遵循相同的配对结构与目录命名规范。 本数据集覆盖多个分类群(taxa),可支持对基因组序列数据进行分析、比较或学习的模型开展跨物种基准测试。数据集旨在支撑多项研究任务,包括序列相似性分析、进化建模、鲁棒性测试,以及针对合成但具备生物学真实性的基因组开展方法评估。 所有序列对均以FASTA文件格式提供,并附带CSV文件,用于标注基因组坐标、重叠群(contigs)、窗口长度以及每一对野生型与合成序列的映射关系。 本数据集为论文《面向真实序列生成的基因组语言模型的基础局限》("Fundamental limitations of genomic language models for realistic sequence generation")的配套数据集,提供了该论文中报告的实验所用的全部原始数据。

提供机构:
Zenodo
创建时间:
2026-01-12
二维码
社区交流群
二维码
科研交流群
商业服务