MNBC-ME identifies mobile elements and putative host species from metagenomic sequences
收藏资源简介:
These files provide supplementary data underlying the article: prok_Next150.fasta.gz: 13441246 150bp-long reads randomly generated from the prokaryotic test genomes, simulating reads sequenced by NextSeq (0.05 coverage) prok_Mi300.fasta.gz: 6723326 300bp-long reads randomly generated from the prokaryotic test genomes, simulating reads sequenced by MiSeq (0.05 coverage) prok_NanoND.fasta.gz: 371870 reads of normally distributed 1kb-10kb lengths randomly generated from the prokaryotic test genomes, simulating reads sequenced by Nanopore (0.05 coverage) plsdb_Next150.fasta.gz: 535249 150bp-long reads randomly generated from the test plasmids, simulating reads sequenced by NextSeq (0.05 coverage) plsdb_Mi300.fasta.gz: 271361 300bp-long reads randomly generated from the test plasmids, simulating reads sequenced by MiSeq (0.05 coverage) plsdb_NanoND.fasta.gz: 24226 reads of normally distributed 1kb-10kb lengths randomly generated from the test plasmids, simulating reads sequenced by Nanopore (0.05 coverage) virus_Next150.fasta.gz: 43381 150bp-long reads randomly generated from the viral test genomes, simulating reads sequenced by NextSeq (0.05 coverage) virus_Mi300.fasta.gz: 23004 300bp-long reads randomly generated from the viral test genomes, simulating reads sequenced by MiSeq (0.05 coverage) virus_NanoND.fasta.gz: 4885 reads of normally distributed 1kb-10kb lengths randomly generated from the viral test genomes, simulating reads sequenced by Nanopore (0.05 coverage) db_list.txt: List of 139052 prokaryotic, plasmidic and viral index filenames in the reference database, indicating the sequence accessions used to build the index files taxonomy.txt: Taxonomy file of the 139052 training sequences. Host taxa are given for plasmids. prok_training_and_test_sequences_list.txt: List of all 56107 prokaryotic training and test sequence accesions. plsdb_training_and_test_sequences_list.txt: List of all 72556 training and test plasmid accesions in the PLSDB database version 2024_05_31_v2 virus_training_and_test_sequences_list.txt: List of all 40353 training and test plasmid accesions in the Virus-Host database release 227
本数据集为对应学术论文提供支撑性补充数据: prok_Next150.fasta.gz:包含13441246条150碱基对(base pair, bp)长度的测序读段,从原核生物测试基因组中随机生成,用于模拟NextSeq测序平台产出的读段(测序覆盖度为0.05) prok_Mi300.fasta.gz:包含6723326条300bp长度的测序读段,从原核生物测试基因组中随机生成,用于模拟MiSeq测序平台产出的读段(测序覆盖度为0.05) prok_NanoND.fasta.gz:包含371870条长度符合正态分布(1千碱基对(kilobase, kb)~10kb)的测序读段,从原核生物测试基因组中随机生成,用于模拟Nanopore测序平台产出的读段(测序覆盖度为0.05) plsdb_Next150.fasta.gz:包含535249条150bp长度的测序读段,从测试质粒中随机生成,用于模拟NextSeq测序平台产出的读段(测序覆盖度为0.05) plsdb_Mi300.fasta.gz:包含271361条300bp长度的测序读段,从测试质粒中随机生成,用于模拟MiSeq测序平台产出的读段(测序覆盖度为0.05) plsdb_NanoND.fasta.gz:包含24226条长度符合正态分布(1kb~10kb)的测序读段,从测试质粒中随机生成,用于模拟Nanopore测序平台产出的读段(测序覆盖度为0.05) virus_Next150.fasta.gz:包含43381条150bp长度的测序读段,从病毒测试基因组中随机生成,用于模拟NextSeq测序平台产出的读段(测序覆盖度为0.05) virus_Mi300.fasta.gz:包含23004条300bp长度的测序读段,从病毒测试基因组中随机生成,用于模拟MiSeq测序平台产出的读段(测序覆盖度为0.05) virus_NanoND.fasta.gz:包含4885条长度符合正态分布(1kb~10kb)的测序读段,从病毒测试基因组中随机生成,用于模拟Nanopore测序平台产出的读段(测序覆盖度为0.05) db_list.txt:参考数据库中139052条原核生物、质粒及病毒的索引文件名列表,标注了用于构建索引文件的序列登录号 taxonomy.txt:139052条训练序列的分类学信息文件,其中质粒序列会标注其宿主分类群 prok_training_and_test_sequences_list.txt:包含全部56107条原核生物训练与测试序列登录号的列表 plsdb_training_and_test_sequences_list.txt:包含PLSDB数据库2024_05_31_v2版本中全部72556条训练与测试质粒序列登录号的列表 virus_training_and_test_sequences_list.txt:包含Virus-Host数据库227版本中全部40353条训练与测试质粒序列登录号的列表



