遇见数据集

Data from: Microhaplotypes provide increased power from short-read DNA sequences for relationship inference

收藏
DataONE2017-11-14 更新2024-06-26 收录
数据链接:
官方服务:

资源简介:

The accelerating rate at which DNA sequence data is now generated by high-throughput sequencing instruments provides both opportunities and challenges for population genetic and ecological investigations of animals and plants. We show here how the common practice of calling genotypes from a single SNP per sequenced region ignores substantial additional information in the phased short-read sequences that are provided by high-throughput sequencing instruments. We target sequenced regions with multiple SNPs in kelp rockfish (Sebastes atrovirens) to determine “microhaplotypes” and then call these microhaplotypes as alleles at each locus. We then demonstrate how these multi-allelic marker data from 96 such loci dramatically increase power for relationship inference. The microhaplotype approach decreases false positive rates by several orders of magnitude, relative to calling bi-allelic SNPs, for two challenging analytical procedures, full sibling and single parent-offspring pair identification. The advent of phased short-read DNA sequence data, in conjunction with emerging analytical tools for their analysis, promises to improve efficiency by reducing the number of loci necessary for a particular level of statistical confidence, thereby lowering the cost of data collection and reducing the degree of physical linkage amongst markers used for relationship estimation. Such advances will facilitate collaborative research and management for migratory and other widespread species.

高通量测序(high-throughput sequencing)设备当下产出DNA序列数据的速率持续加快,这为动植物的种群遗传学与生态学研究带来了机遇与挑战。我们在此展示,当前通行的在每个测序区域仅基于单个单核苷酸多态性(Single Nucleotide Polymorphism,简称SNP)进行基因型分型的常规操作,会忽略高通量测序设备产出的相位化短读长序列中蕴含的大量额外信息。我们以海带平鲉(Sebastes atrovirens)中携带多个SNP的测序区域为研究对象,确定其"微单倍型(microhaplotype)",随后将这些微单倍型作为每个基因座(locus)的等位基因进行分型。随后我们证明,来自96个此类基因座的多等位基因标记数据,可大幅提升亲缘关系推断的效能。相较于仅对双等位基因SNP进行分型的方法,在全同胞及单亲-子代配对鉴定这两项颇具挑战性的分析流程中,微单倍型策略可将假阳性率降低数个数量级。相位化短读长DNA序列数据的出现,结合新兴的配套分析工具,有望通过缩减达成特定统计置信度所需的基因座数量以提升研究效率,此举可降低数据采集成本,同时减少亲缘关系推断所用标记间的物理连锁程度。此类技术进展将助力洄游物种及其他广布物种的协同研究与管理工作。

创建时间:
2017-11-14
二维码
社区交流群
二维码
科研交流群
商业服务