Vantage-point search (vpsearch) dataset and index
收藏资源简介:
This is a dataset for use with the vpsearch tool. This dataset contains a compressed sequence file, bac120_ssu_reps_r207-sliced-dedup.fa.gz, which was obtained from the GTDB dataset of bacterial sequences (v207) by extracting the v3-v4 hypervariable region and removing duplicate sequences. Also included is a vantage-point tree, built with version 0.1.2 of the vpsearch software. Note: the vantage-point tree needs to be decompressed (tar xzvf bac120_ssu_reps_r207-sliced-dedup.db.tar.gz) before it can be used. The scripts used to prepare the dataset can be found in the vpsearch GitHub repository.
本数据集专为vpsearch工具设计使用。数据集包含一份压缩序列文件`bac120_ssu_reps_r207-sliced-dedup.fa.gz`,该文件源自v207版本的基因组分类数据库(Genome Taxonomy Database,GTDB)细菌序列数据集,经提取V3-V4高变区并移除重复序列后得到。此外还包含使用vpsearch软件v0.1.2版本构建的视点树(vantage-point tree)。注意:该视点树需先执行解压命令`tar xzvf bac120_ssu_reps_r207-sliced-dedup.db.tar.gz`后方可使用。用于制备本数据集的脚本可在vpsearch的GitHub仓库中获取。



