遇见数据集

bolinas-dna/zoonomia-v1-v3_utr3

收藏
Hugging Face2026-05-11 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是`bolinas-dna/zoonomia-v1-v1`的一个分区子集,专门针对标记为`utr3`(3非翻译区)的锚点区域。`utr3`区域来源于Ensembl r115的蛋白质编码转录本的3 UTR,并过滤为仅包含蛋白质编码类型。在区域标签优先级划分中,`utr3`的优先级仅次于`cds`,高于其他非编码区域。该子集包含54,828个人类锚点(占原始v1数据集的4.82%),通过halLiftover投影扩展到108种Zoonomia哺乳动物,并经过反向互补增强,最终形成10,305,306个训练样本。数据模式与原始v1相同,包括序列、物种、染色体位置和增强信息等列。该数据集用于基因组学和生物学研究,专注于跨哺乳动物的3非翻译区比较分析。

This dataset is a partitioned subset of `bolinas-dna/zoonomia-v1-v1`, specifically focused on anchor regions labeled `utr3` (3 untranslated region). The `utr3` regions are derived from the 3 UTR of Ensembl r115 protein-coding transcripts, filtered to include only protein-coding biotypes. In the region label priority hierarchy, `utr3` has the second-highest priority after `cds`, overriding other non-coding regions. The subset contains 54,828 human anchors (4.82% of the original v1 dataset), expanded to 10,305,306 training samples via halLiftover projection to 108 Zoonomia mammals and reverse-complement augmentation. The schema matches the original v1, including columns such as sequence, species, chromosomal positions, and augmentation. This dataset is designed for genomics and biology research, focusing on comparative analysis of 3 untranslated regions across mammals.

提供机构:
bolinas-dna
二维码
社区交流群
二维码
科研交流群
商业服务