Data from: Identifying and reducing AFLP genotyping error: an example of tradeoffs when comparing population structure in broadcast spawning versus brooding oysters
收藏资源简介:
Phylogeographic inferences about gene flow are strengthened through comparison of co-distributed taxa, but also depend on adequate genomic sampling. Amplified Fragment Length Polymorphisms (AFLP) provide a rapid and inexpensive source of multilocus allele frequency data for making genomically robust inferences. Every AFLP study initially generates markers with a range of locus-specific genotyping error rates and applies criteria to select a subset for analysis. However, there has been very little empirical evaluation of the best tradeoff between culling all but the lowest-error loci to minimize overall genotyping error versus the potential for increasing population genetic signal by retaining more loci. Here, we used AFLPs to compare population structure in co-distributed broadcast spawning (Crassostrea virginica) and brooding (Ostrea equestris) oyster species. Using existing methods for almost entirely automated marker selection and scoring, genotyping error tradeoffs were evaluated by comparing results across a nested series of datasets with mean mismatch errors of 0, 1, 2, 3, 4 and >4%. Artifactual population structure was diagnosed in high-error datasets and we assessed the low-error point at which expected population substructure signal was lost. In both species we identified substructure patterns deemed to be inaccurate at error rates {less than or equal to}2% and >4%. In the species comparison, the optimum datasets showed higher gene flow for the brooding oyster with more oceanic salinity tolerances. AFLP tradeoffs may differ among studies, but our results suggest that important signal may be lost in the pursuit of 'acceptable' error levels and our procedures provide a general method for empirically exploring these tradeoffs.
关于基因流的系统地理推断,可通过对同域分布类群的比较得到强化,但同时也依赖于充足的基因组采样。扩增片段长度多态性(Amplified Fragment Length Polymorphisms, AFLP)能够快速且低成本地获取多位点等位基因频率数据,用于开展具有基因组可靠性的推断。所有AFLP研究最初都会生成一系列具有不同位点特异性基因分型错误率的分子标记,并通过设定筛选标准选取部分标记用于后续分析。然而,目前鲜有实证研究对二者的最优权衡展开评估:一方面是剔除除低错误率位点外的所有位点以最小化整体基因分型错误,另一方面是保留更多位点以提升种群遗传信号强度。 本研究利用AFLP技术,对同域分布的体外产卵(Broadcast Spawning)型(Crassostrea virginica)与育幼(Brooding)型(Ostrea equestris)两种牡蛎的种群结构开展比较分析。本研究采用近乎全自动化的标记筛选与打分方法,通过对比平均错配错误率分别为0、1、2、3、4以及>4%的一系列嵌套数据集的分析结果,对基因分型错误的权衡效应进行评估。在高错误率数据集内发现了人为假象的种群结构,同时本研究还测定了预期种群亚结构信号消失的最低错误率阈值。 在两种牡蛎中,我们均发现当错误率≤2%或>4%时,所检测到的亚结构模式被判定为不准确。 跨物种比较结果显示,最优数据集表明育幼型牡蛎(其海洋盐度耐受性更强)的基因流水平更高。 不同研究中AFLP分析的权衡策略或存在差异,但本研究结果表明,一味追求‘可接受’的错误率阈值可能会丢失重要的种群遗传信号,而本研究提出的流程可为实证探究这类权衡效应提供通用方法。



