遇见数据集

Refseq Test Subsets for Frame Classification with and without Errors

收藏
NIAID Data Ecosystem2026-03-12 收录
数据链接:
官方服务:

资源简介:

These test files extend the 'Refseq datasets for training frame classification' dataset. It provided the original test file and three variations containing erroneous sequences to simulate realistic data. The data is based on randomly selected viral and bacterial genomes and the human193(GRCh38.p13) reference genome which was downloaded from GenBank. From each original nucleic acid sequence, we created multiple patches of length 300 in all possible reading frames using a sliding window on the initial sequence and its reversed complement. The data is stored in the FASTA format according to the following convention: >{ID}_subsequence{patch index}_frame{frame index}|{class marker}|{frame index} sequence with ID - denotes the ReSeq accession of the original sequence in the Refseq dataset. sequence - nucleic acid sequence patch of length 300 or 250 patch_index - denotes the starting triplet of the given patch within the original sequence or reverse complemented sequence (i.e. 3*patch_index is the starting index of frame 0 in the original sequence) class marker - indicates the taxonomic domain 0 - virus 1 - bacteria 2 - human / mammal frame index - indicates the reading frame 0 - on-frame 1 - shifted by one 2 - shifted by two 3 - reverse complemented 4 - shifted by one and reverse complemented 5 - shifted by two and reverse complemented Each file contains 212.618 patches per frame.

本测试文件集拓展了用于帧分类训练的RefSeq数据集(Refseq datasets for training frame classification)。该数据集包含原始测试文件与三份含错误序列的变体文件,用以模拟真实场景下的数据。 本数据集的数据来源于随机选取的病毒、细菌基因组,以及从基因序列数据库(GenBank)下载的人类193(GRCh38.p13)参考基因组。研究团队针对每条原始核酸序列,利用滑动窗口在初始序列及其反向互补序列上进行截取,生成所有可能可读框下的多段长度为300的序列片段(patch)。 本数据集的数据按照以下规范以FASTA格式存储: >{ID}_subsequence{patch index}_frame{frame index}|{class marker}|{frame index} sequence 各字段说明如下: - ID:指代RefSeq数据集中原始序列的ReSeq登录号(ReSeq accession) - sequence:长度为300或250的核酸序列片段(patch) - patch_index:表示当前片段在原始序列或反向互补序列中的起始三联体位置(即3*patch_index为原始序列中可读框0的起始索引) - class marker:表示所属分类域 - 0:病毒 - 1:细菌 - 2:人类/哺乳动物 - frame index:表示可读框类型 - 0:正链可读框 - 1:偏移1个碱基的可读框 - 2:偏移2个碱基的可读框 - 3:反向互补序列可读框 - 4:偏移1个碱基且反向互补的可读框 - 5:偏移2个碱基且反向互补的可读框 每个文件的每个可读框均包含212618个序列片段。

创建时间:
2021-10-05
二维码
社区交流群
二维码
科研交流群
商业服务