bolinas-dna/zoonomia-v1-v3_tss_region_and_utr5
收藏资源简介:
该数据集名为`bolinas-dna/zoonomia-v1-v3_tss_region_and_utr5`,是一个生物学和基因组学领域的DNA数据集,专门针对转录起始位点(TSS)区域和5非翻译区(UTR)。它是从跨哺乳动物训练集`bolinas-dna/zoonomia-v1-v1`中按区域类型分区得到的,仅包含标记为`tss_region_and_utr5`的锚点区域。该区域标签覆盖Ensembl r115转录本的TSS±256 bp区域以及所有蛋白质编码转录本的5 UTR,作为一个类别处理,因为启动子和5 UTR在构造上重叠。数据集是v1数据集的分区之一,通过优先级划分方法从六个v3子集中选择,包含40,823个人类锚点(占v1的3.59%),通过halLiftover投影到108种Zoonomia哺乳动物并进行反向互补增强后,扩展为8,124,514个训练样本。数据模式与原始v1数据集相同,包括查询名称、物种、染色体位置、序列和增强信息等列。构建过程包括构建v1训练集、注释区域标签、过滤锚点、增强和分片。注意事项指出v3子集是v1的分区而非独立探针,背景区域具有异质性,且`ncrna_exon`标签定义较宽。
The dataset `bolinas-dna/zoonomia-v1-v3_tss_region_and_utr5` is a per-anchor region-type partition of the cross-mammal training set `bolinas-dna/zoonomia-v1-v1`, restricted to anchors labelled `tss_region_and_utr5` by the `snakemake/zoonomia_projection_dataset` pipeline. The region label `tss_region_and_utr5` represents TSS region and 5 UTR — (TSS ± 256 bp on every Ensembl r115 transcript) ∪ 5 UTR of every protein-coding transcript, treated as one class due to overlap between promoters and 5 UTRs. It is one of six v3 subsets partitioning the conservation-filtered human anchor set by priority-walk, containing 40,823 of 1,136,854 human anchors (3.59% of v1), expanding to 8,124,514 training samples after halLiftover projection to up to 108 Zoonomia mammals and reverse-complement augmentation. The schema matches the original v1 dataset, with columns such as `query_name`, `species`, `t_chrom`, `t_start`, `t_end`, `t_strand`, `t_src_size`, `sequence`, and `augmentation`. Construction involves building the v1 set, annotating region labels, filtering anchors, RC-augmenting, shuffling, and sharding. Caveats include that the v3 subsets are a partition of v1, not independent probes, and background regions are heterogeneous.



