Binned Nanopore/Illumina Reads From DNAformer, Split by the Data Type
收藏资源简介:
This dataset includes binned DNA sequencing reads from Nanopore and Illumina platforms, used in DNA data storage experiments (DNAformer). Each file contains read clusters (label, separator of "****", reads) separated by two blank lines. Files are split based on the file type: _Random.txt: clusters of the random file. _Semantic.txt: clusters of the semantic file. Includes data from two Nanopore flowcells (separate and merged) and a test Illumina set. Additionally, the semantic file and the random file are included in the repository (see semantic_file.zip and rand_file.bin), as well and their corresponding encoded DNA sequences are given in sequences_semantic_file.txt and sequences_random_file.txt.
本数据集包含源自纳米孔(Nanopore)与Illumina测序平台的分箱DNA测序读段(reads),用于DNA数据存储实验(DNAformer)。每个文件内的测序读段簇由标签、分隔符“****”与测序读段(reads)构成,簇间以两行空行分隔。 数据集按文件类型划分为: _Random.txt:对应随机文件的读段簇。 _Semantic.txt:对应语义文件的读段簇。 本数据集包含两组独立及合并后的纳米孔测序流动槽(flowcell)数据,以及一套Illumina测试数据集。 此外,本仓库还收录了该语义文件与随机文件(详见semantic_file.zip与rand_file.bin),其对应的编码DNA序列分别存储于sequences_semantic_file.txt与sequences_random_file.txt中。



