Indexing and searching petabyte-scale nucleotide resources
收藏资源简介:
Searching vast and rapidly growing nucleotide content in resources, such as runs in Sequence Read Archive and assemblies for whole genome shotgun sequencing projects in GenBank, is currently impractical for most researchers. We present Pebblescout, a tool that navigates such content by providing indexing and search capabilities. Indexing uses dense sampling of the sequences in the resource. Search finds subjects (runs or assemblies) that have short sequence matches to a user query, with well-defined guarantees, and ranks them using informativeness of the matches. Seven databases that index over 3.5 petabases were created to illustrate the functionality of Pebblescout. They can be interactively searched at https://pebblescout.ncbi.nlm.nih.gov. Here we show that for a wide range of query lengths, Pebblescout provides a data-driven way for finding relevant subsets of large nucleotide resources. We compare Pebblescout to MetaGraph and Sourmash and show that both methods do not report some results correctly reported by Pebblescout.
对于多数研究人员而言,检索诸如序列读取存档(Sequence Read Archive, SRA)中的测序运行数据、以及基因银行(GenBank)内全基因组鸟枪测序项目的组装序列这类规模庞大且持续增长的核苷酸序列资源,目前尚不具备可行性。为此我们推出Pebblescout,一款通过构建索引与提供检索功能来实现这类资源导航的工具。该工具的索引构建过程会对资源内的序列进行密集采样;检索环节则可定位与用户查询序列存在短片段匹配的目标(测序运行或组装序列),并附带明确的可靠性保障,同时依据匹配片段的信息价值对目标进行排序。为展示Pebblescout的功能,我们构建了7个总索引规模超过3.5拍碱基的数据库,用户可通过https://pebblescout.ncbi.nlm.nih.gov进行交互式检索。本文证明,在广泛的查询序列长度范围内,Pebblescout均可为研究者提供一种数据驱动的方式,以从大规模核苷酸资源中筛选出相关子集。我们将Pebblescout与MetaGraph、Sourmash进行了对比,结果显示后两种方法均存在部分Pebblescout可正确检索到的结果漏报情况。



