Data from: On the accuracy of genomic selection
收藏资源简介:
Genomic selection is focused on prediction of breeding values of selection candidates by means of high density of markers. It relies on the assumption that all quantitative trait loci (QTLs) tend to be in strong linkage disequilibrium (LD) with at least one marker. In this context, we present theoretical results regarding the accuracy of genomic selection, i.e., the correlation between predicted and true breeding values. Typically, for individuals (so-called test individuals), breeding values are predicted by means of markers, using marker effects estimated by fitting a ridge regression model to a set of training individuals. We present a theoretical expression for the accuracy; this expression is suitable for any configurations of LD between QTLs and markers. We also introduce a new accuracy proxy that is free of the QTL parameters and easily computable; it outperforms the proxies suggested in the literature, in particular, those based on an estimated effective number of independent loci (Me). The theoretical formula, the new proxy, and existing proxies were compared for simulated data, and the results point to the validity of our approach. The calculations were also illustrated on a new perennial ryegrass set (367 individuals) genotyped for 24,957 single nucleotide polymorphisms (SNPs). In this case, most of the proxies studied yielded similar results because of the lack of markers for coverage of the entire genome (2.7 Gb).
基因组选择(Genomic selection)旨在借助高密度标记,对候选选系的育种值进行预测。该方法基于如下假设:所有数量性状基因座(quantitative trait loci, QTLs)均倾向于与至少一种标记存在较强的连锁不平衡(linkage disequilibrium, LD)。 于此背景下,本文针对基因组选择的准确度——即预测育种值与真实育种值间的相关系数——给出了理论推导结果。通常而言,针对个体(即所谓测验个体)的育种值预测,可借助标记完成:先通过岭回归模型对训练群体的个体进行拟合以估计标记效应,再基于该效应实现预测。 本文推导了适用于任意数量性状基因座与标记间连锁不平衡构型的准确度理论表达式;同时提出了一种无需数量性状基因座参数、且易于计算的新型准确度代理指标,该指标的性能优于现有文献中提出的各类代理指标,尤其是基于独立基因座有效数目(effective number of independent loci, Me)的代理指标。 通过模拟数据对本文提出的理论公式、新型代理指标及现有代理指标进行了对比验证,结果证实了本方法的有效性。此外,本文还以一组新的多年生黑麦草群体(共367个个体)为例进行了演示:该群体经24957个单核苷酸多态性(single nucleotide polymorphisms, SNPs)分型,由于缺乏足够标记以覆盖2.7 Gb的全基因组,本研究中多数代理指标的预测结果趋于一致。



