遇见数据集

vedatonuryilmaz/te-seqdata-v1

收藏
Hugging Face2026-05-21 更新2026-05-31 收录
官方服务:

资源简介:

TE-Seq Data v1是一个人类基因组(GRCh38)转座因子位点序列数据集,包含4,693,511条记录,文件大小约为894 MB。数据集中每条记录包含DNA序列(以原始形式提供,无填充或修剪)、转座因子家族标签(作为主要训练目标,有1,180个唯一值)、转座因子类别标签(如SINE、LINE、LTR、DNA等,有13个唯一值),以及基因组坐标(染色体、起始位置、结束位置)、链方向(+或-)、亚家族标签和其他元数据列(如ID、位点信息等)。该数据集设计用于DeepGenopix工具的训练、模拟和基准测试任务,提供了修正后的分割文件(训练集、验证集、测试集),并自动基于样本数量进行分层划分。类别分布显示,SINE、LINE、LTR和DNA是主要类别,数量分别为1,770,903、1,516,226、720,177和483,994条。

TE-Seq Data v1 is a dataset of transposable element locus sequences from the human genome (GRCh38), containing 4,693,511 records with a file size of approximately 894 MB. Each record includes a DNA sequence (to be consumed verbatim without padding or trimming), a TE family label (the primary training target with 1,180 unique values), a TE class label (e.g., SINE, LINE, LTR, DNA, with 13 unique values), genomic coordinates (chromosome, start, end), strand orientation (+ or -), subfamily label, and other metadata columns (such as ID, loci, group, etc.). The dataset is intended for use in DeepGenopix training, simulator, and benchmark jobs, with corrected split files (train, validation, test) that are automatically stratified based on sample counts. The class distribution shows SINE, LINE, LTR, and DNA as the major classes, with counts of 1,770,903, 1,516,226, 720,177, and 483,994 respectively.

提供机构:
vedatonuryilmaz
二维码
社区交流群
二维码
科研交流群
商业服务