bolinas-dna/zoonomia-v1-v3_ncrna_exon
收藏资源简介:
`bolinas-dna/zoonomia-v1-v3_ncrna_exon` 是一个基于跨哺乳动物训练集的数据集分区,专注于非编码RNA外显子区域。它是从`bolinas-dna/zoonomia-v1-v1`数据集中按锚点区域类型分区得到的,限制为通过`snakemake/zoonomia_projection_dataset`管道标记为`ncrna_exon`的锚点。区域标签`ncrna_exon`定义为非编码RNA外显子,包括Ensembl r115中所有不属于蛋白质编码转录本的外显子,没有生物类型或质量过滤,因此比`val_ncrna`验证配方更广泛。在分区中,通过优先级遍历将人类锚点集分为六个v3子集,本子集包含93,064个锚点(占v1的8.19%),通过投影到108个Zoonomia哺乳动物和反向互补增强后扩展到15,277,064个训练样本。数据模式与原始v1相同,包括查询名称、物种、序列等列。构造过程涉及构建v1训练集、标注区域标签、过滤到`ncrna_exon`、RC增强和分片。注意事项包括v3子集是v1的分区而非独立探针、`ncrna_exon`定义较宽、背景区域异质性等。
`bolinas-dna/zoonomia-v1-v3_ncrna_exon` is a per-anchor region-type partition of the cross-mammal training set `bolinas-dna/zoonomia-v1-v1`, restricted to anchors labelled `ncrna_exon` by the `snakemake/zoonomia_projection_dataset` pipeline. The region label `ncrna_exon` refers to non-coding-RNA exons, defined as every Ensembl r115 exon that is not part of a protein-coding transcript, with no biotype or quality filter, making it broader than the `val_ncrna` validation recipe. The partition divides the human anchor set into six v3 subsets via priority-walk, and this subset contains 93,064 anchors (8.19% of v1), expanding to 15,277,064 training samples after projection to up to 108 Zoonomia mammals and reverse-complement augmentation. The schema matches the original v1 dataset, with columns like query_name, species, sequence, etc. Construction involves building the v1 training set, annotating anchors with region labels, filtering to `ncrna_exon`, RC augmentation, and sharding. Caveats include that the v3 subsets are a partition of v1, not independent probes; `ncrna_exon` is broadly defined; and background regions are heterogeneous.



