遇见数据集

Data from: Fast and cost-effective genetic mapping in apple using next-generation sequencing

收藏
DataONE2014-07-22 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

Next-generation DNA sequencing (NGS) produces vast amounts of DNA sequence data, but it is not specifically designed to generate data suitable for genetic mapping. Recently developed DNA library preparation methods for NGS have helped solve this problem, however, by combining the use of reduced representation libraries with DNA sample barcoding to generate genome-wide genotype data from a common set of genetic markers across a large number of samples. Here we use such a method, called genotyping-by-sequencing (GBS), to produce a data set for genetic mapping in an F1 population of apples (Malus x domestica) segregating for skin color. We show that GBS produces a relatively large, but extremely sparse, genotype matrix: over 270,000 SNPs were discovered, but most SNPs have too much missing data across samples to be useful for genetic mapping. After filtering for genotype quality and missing data, only 6% of the 85 million DNA sequence reads contributed to useful genotype calls. Despite this limitation, using existing software and a set of simple heuristics, we generated a final genotype matrix containing 3967 SNPs from 89 DNA samples from a single lane of Illumina HiSeq and used it to create a saturated genetic linkage map and to identify a known QTL underlying apple skin color. We therefore demonstrate that GBS is a cost effective method for generating genome-wide SNP data suitable for genetic mapping in a highly diverse and heterozygous agricultural species. We anticipate future improvements to the GBS analysis pipeline presented here that will enhance the utility of next-generation DNA sequence data for the purposes of genetic mapping across diverse species.

新一代DNA测序(Next-generation DNA sequencing, NGS)可产生海量DNA序列数据,但该技术并非专为生成适用于遗传作图的数据而设计。不过,近期开发的NGS用DNA文库制备方法通过将简化基因组文库与DNA样本条码标记相结合,可从大量样本的一套通用遗传标记中生成全基因组基因型数据,从而解决了这一难题。本研究采用一种被称为测序分型(genotyping-by-sequencing, GBS)的方法,为苹果(Malus x domestica)的一个因果皮颜色而发生分离的F1群体构建遗传作图用数据集。研究表明,GBS可生成规模相对较大但极为稀疏的基因型矩阵:共发现超过27万个单核苷酸多态性(Single Nucleotide Polymorphism, SNPs),但大多数SNPs在各样本中存在过多缺失数据,无法用于遗传作图。在对基因型质量和缺失数据进行过滤后,8500万条DNA测序读段中仅有6%可用于生成有效的基因型分型结果。尽管存在这一局限,本研究借助现有软件与一组简单启发式算法,利用单张Illumina HiSeq测序通道上的89个DNA样本,生成了包含3967个SNPs的最终基因型矩阵,并以此构建了饱和遗传连锁图谱,同时鉴定出一个控制苹果果皮颜色的已知数量性状基因座(Quantitative Trait Locus, QTL)。因此,本研究证明,GBS是一种经济高效的方法,可用于在高度多样化且杂合的农业物种中生成适用于遗传作图的全基因组SNP数据。本研究期待对本文所提出的GBS分析流程进行未来改进,以提升新一代DNA测序数据在不同物种遗传作图中的应用价值。

创建时间:
2014-07-22
二维码
社区交流群
二维码
科研交流群
商业服务