遇见数据集

bolinas-dna/zoonomia-v1-v3_cds

收藏
Hugging Face2026-05-11 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是bolinas-dna/zoonomia-v1-v1数据集的子集,专门针对编码序列(CDS)区域。它通过优先级划分从跨哺乳动物训练集中筛选出402,393个人类锚点(占原数据集的35.40%),并扩展到78,387,828个训练样本,覆盖108种Zoonomia哺乳动物。数据包含查询名称、物种、染色体位置、序列(255 bp)和增强信息等字段,格式为JSONL.zst分片。数据集是六个v3子集之一,这些子集共同构成完整数据集的分区,每个锚点仅出现在一个子集中。

This dataset is a subset of bolinas-dna/zoonomia-v1-v1, specifically for coding sequence (CDS) regions. It partitions the cross-mammal training set by priority, containing 402,393 human anchors (35.40% of v1) and expanding to 78,387,828 training samples across 108 Zoonomia mammals. The data includes query name, species, chromosome location, sequence (255 bp), and augmentation information, stored in JSONL.zst shards. It is one of six v3 subsets that form a partition of the full dataset, with each anchor appearing in exactly one subset.

提供机构:
bolinas-dna
二维码
社区交流群
二维码
科研交流群
商业服务