遇见数据集

Phylo-k-mers databases for SHERPAS

收藏
Dryad2024-12-29 收录
官方服务:

资源简介:

SHERPAS is a new program to identify novel recombinant sequences in a large collection of viral sequences, and to provide a first estimate of their recombinant structure. SHERPAS is much faster than other softwares for recombination detection; its main feature is the use of a pre-computed database of "phylogenetically-informed k-mers" (or phylo-k-mers). The computation of this phylo-k-mer database is a heavy computational step, but it only needs to be executed once for a given reference alignment. A phylo-k-mer database can be built from any reference alignment, and a phylogenetic tree built from that alignment, using RAPPAS2 (https://github.com/phylo42/rappas2). We propose here three ready-to-use databases, for three reference alignments: -An alignment of 167 sequences of the pol region of the HIV genome, provided with the program SCUEAL, accessible at https://github.com/spond/SCUEAL/blob/master/data/pol2009.nex -An alignment of 339 sequence of the whole HBV genome, provided with the programm jpHMM, accessible at http://jphmm.gobics.de/download.html. -An alignment of 881 sequences of the whole HIV genome, also provided with jpHMM, accessible at http://jphmm.gobics.de/download.html. For each of these alignments, we provide a .zip file containing three files: The phylo-k-mer database (.rps file), the reference phylogenetic tree used to build the database (.tree file), and a table associating each reference sequence to a strain of the virus (.csv file). The details of the construction of the database, the construction of the tree, as well as the origin of the information reported in the table, can be found in the Supplementary Materials associated with the original Bioinformatics publication.

SHERPAS是一款可在大规模病毒序列集合中识别新型重组序列,并对其重组结构进行首次估算的全新程序。相较于其他重组检测软件,SHERPAS的运行速度大幅提升,其核心特性在于采用了预计算的「系统发育信息型k-mer(phylogenetically-informed k-mers,简称phylo-k-mers)」数据库。该phylo-k-mer数据库的计算过程属于计算密集型步骤,但针对特定的参考序列比对仅需执行一次。phylo-k-mer数据库可通过RAPPAS2(https://github.com/phylo42/rappas2),由任意参考序列比对以及基于该比对构建的系统发育树来构建。 本次共提供三款适配不同参考比对的预构建可用数据库: 1. 适配SCUEAL程序附带的167条HIV基因组pol区序列比对文件,该文件可在https://github.com/spond/SCUEAL/blob/master/data/pol2009.nex获取; 2. 适配jpHMM程序附带的339条全HBV基因组序列比对文件,可在http://jphmm.gobics.de/download.html下载; 3. 同样由jpHMM程序附带的881条全HIV基因组序列比对文件,亦可在http://jphmm.gobics.de/download.html获取。 针对上述每一组参考比对,我们均提供了一个.zip压缩包,内含三类文件:用于构建数据库的phylo-k-mer数据库文件(.rps格式)、用于构建该数据库的参考系统发育树文件(.tree格式),以及关联每条参考序列与病毒毒株的对照表(.csv格式)。 关于该数据库的构建流程、系统发育树的构建方法,以及对照表中信息的来源详情,均可查阅原发表于《Bioinformatics》(生物信息学)期刊的研究论文的补充材料。

二维码
社区交流群
二维码
科研交流群
商业服务