Vantage-point search (vpsearch) dataset and index
收藏资源简介:
This is a dataset for use with the vpsearch tool. This dataset contains the following files: A compressed sequence file, bac120_ssu_reps_r207-sliced-dedup.fa.gz, which was obtained from the GTDB dataset of bacterial sequences (v207) by extracting the v3-v4 hypervariable region and removing duplicate sequences. A vantage-point tree, built with version 0.1.2 of the vpsearch software. Note: the vantage-point tree needs to be decompressed (tar xzvf bac120_ssu_reps_r207-sliced-dedup.db.tar.gz) before it can be used for querying. A sample query file, query.fa, containing the sequence NR_126253.1 from RefSeq. The scripts used to prepare the dataset can be found in the vpsearch GitHub repository. Also included in the repository is a description of how the primary data was obtained.



