Data from: On the accuracy of genomic selection
收藏资源简介:
Genomic selection is focused on prediction of breeding values of selection candidates by means of high density of markers. It relies on the assumption that all quantitative trait loci (QTLs) tend to be in strong linkage disequilibrium (LD) with at least one marker. In this context, we present theoretical results regarding the accuracy of genomic selection, i.e., the correlation between predicted and true breeding values. Typically, for individuals (so-called test individuals), breeding values are predicted by means of markers, using marker effects estimated by fitting a ridge regression model to a set of training individuals. We present a theoretical expression for the accuracy; this expression is suitable for any configurations of LD between QTLs and markers. We also introduce a new accuracy proxy that is free of the QTL parameters and easily computable; it outperforms the proxies suggested in the literature, in particular, those based on an estimated effective number of independent loci (Me). The theoretical formula, the new proxy, and existing proxies were compared for simulated data, and the results point to the validity of our approach. The calculations were also illustrated on a new perennial ryegrass set (367 individuals) genotyped for 24,957 single nucleotide polymorphisms (SNPs). In this case, most of the proxies studied yielded similar results because of the lack of markers for coverage of the entire genome (2.7 Gb).
基因组选择(Genomic selection)旨在借助高密度分子标记,对候选育种个体的育种值进行预测。该方法的核心假设为:所有数量性状位点(Quantitative Trait Loci, QTL)均至少与一个分子标记存在较强的连锁不平衡(Linkage Disequilibrium, LD)。基于此背景,本文针对基因组选择的预测准确性——即预测育种值与真实育种值之间的相关系数——给出了理论推导结果。常规研究流程中,针对待测个体(即所谓的测试个体),育种值的预测依托分子标记完成:先通过岭回归模型对训练群体的个体进行拟合以估计标记效应,再基于该效应完成待测个体育种值的预测。本文推导了适用于数量性状位点与标记间任意连锁不平衡构型的准确性理论表达式;同时提出了一种无需数量性状位点参数、易于计算的新型准确性代理指标,该指标优于现有文献中提出的各类代理指标,尤其是基于估计有效独立位点数目(Me)的代理指标。本研究通过模拟数据对理论公式、新型代理指标及现有代理指标进行了对比验证,结果证实了本文方法的有效性。此外,本研究以一个新型多年生黑麦草群体(共367个个体,对24957个单核苷酸多态性(Single Nucleotide Polymorphism, SNPs)进行了基因分型)为例,对计算流程进行了演示。本次分析中,由于该群体的分子标记覆盖率不足(基因组总大小为2.7 Gb),多数被研究的代理指标结果趋于一致。



