遇见数据集

Data from: Not all sequence tags are created equal: designing and validating sequence identification tags robust to indels

收藏
DataONE2012-08-13 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

Ligating adapters with unique synthetic oligonucleotide sequences (sequence tags) onto individual DNA samples before massively parallel sequencing is a popular and efficient way to obtain sequence data from many individual samples. Tag sequences should be numerous and sufficiently different to ensure sequencing, replication, and oligonucleotide synthesis errors do not cause tags to be unrecoverable or confused. However, many design approaches only protect against substitution errors during sequencing and extant tag sets contain too few tag sequences. We developed an open-source software package to validate sequence tags for conformance to two distance metrics and design sequence tags robust to indel and substitution errors. We use this software package to evaluate several commercial and non-commercial sequence tag sets, design several large sets (maxcount=7,198) of edit metric sequence tags having different lengths and degrees of error correction, and integrate a subset of these edit metric tags to polymerase chain reaction (PCR) primers and sequencing adapters. We validate a subset of these edit metric tagged PCR primers and sequencing adapters by sequencing on several platforms and subsequent comparison to commercially available alternatives. We find that several commonly used sets of sequence tags or design methodologies used to produce sequence tags do not meet the minimum expectations of their underlying distance metric, and we find that PCR primers and sequencing adapters incorporating edit metric sequence tags designed by our software package perform as well as their commercial counterparts. We suggest that researchers evaluate sequence tags prior to use or evaluate tags that they have been using. The sequence tag sets we design improve on extant sets because they are large, valid across the set, and robust to the suite of substitution, insertion, and deletion errors affecting massively parallel sequencing workflows on all currently used platforms.

在大规模并行测序(massively parallel sequencing)前,将带有独特合成寡核苷酸(oligonucleotide)序列的连接接头连接至单个DNA样本,是从众多个体样本中获取序列数据的一种高效通用手段。序列标签(sequence tags)需数量充足且差异度足够,以避免测序、复制及寡核苷酸合成过程中出现的错误导致标签无法识别或混淆。然而,多数现有设计方法仅能抵御测序过程中的替换错误,且现存的序列标签集合的标签数量严重不足。我们开发了一款开源软件包,用于验证序列标签是否符合两种距离度量标准,并设计出可抵御插入缺失(indel)与替换错误的序列标签。我们利用该软件包对多款商用及非商用序列标签集合进行评估,设计出多组规模庞大(最大数量maxcount=7198)、具备不同长度与纠错等级的编辑距离度量(edit metric)序列标签集合,并将其中部分编辑距离度量标签整合至聚合酶链式反应(polymerase chain reaction, PCR)引物与测序接头中。我们通过多个测序平台对部分搭载编辑距离度量标签的PCR引物与测序接头进行测序验证,并将结果与市售同类产品进行对比。我们发现,多款常用序列标签集合或其制备所用的设计方法均未达到其底层距离度量标准的最低要求;同时,搭载本软件包设计的编辑距离度量序列标签的PCR引物与测序接头,其性能可与市售同类产品媲美。我们建议研究人员在使用序列标签前对其进行评估,或对自身已在使用的标签开展评估工作。我们所设计的序列标签集合相较现存集合实现了优化:其规模庞大、集合内标签均符合校验标准,且可抵御当前所有主流测序平台的大规模并行测序流程中可能出现的各类替换、插入与缺失错误。

创建时间:
2012-08-13
二维码
社区交流群
二维码
科研交流群
商业服务