Prediction Pipeline for Optimizing resource allocation in Miscanthus breeding with sparse testing designs for genomic prediction
收藏资源简介:
Phenotyping high-biomass perennial crops is laborious and the rate of genetic gain in perennial crop breeding programs is typically low. So, it is especially important to identify methods that produce efficiency gains in the breeding process. Miscanthus is a C4 perennial grass with favorable characteristics for producing biomass as a feedstock for biofuels and diverse biobased products. Increasing biomass yield will increase profitability and environmental benefits, so is a key target for Miscanthus breeding. In addition, the identification of well-adapted genotypes across a wide range of environmental conditions requires the establishment of multi-environment trials (METs). Sparse testing is a genomic prediction-based strategy that reduces the phenotyping costs in METs by selecting a subset of genotypes to evaluate in a subset of environments and then predicts the performance of the unobserved genotype-environment combinations. A Miscanthus sacchariflorus (MSA) population comprising 336 genotypes observed across three environments was analyzed. Three prediction models considering main effects (environments, genotypes, genomic) and interaction effects (genotype-by-environment; G×E interaction) were implemented for forecasting dry biomass yield (YDY), total culm (TCM), average internode length (AIL), and culm node number (CNN). Multiple calibration sets based on different compositions and sizes were considered to evaluate performance in terms of the predictive ability (PA) and the mean square error (MSE) for a fixed testing set size. The training set size ranged from 52 to 112 to predict a fixed set of 224 unobserved genotypes across all three environments. The results showed that the model accounting for G×E interaction presented the highest PA and the lowest MSE for CNN (PA: ~0.77, MSE: ~0.5) and YDY (PA: ~0.70, MSE: ~1.3) while for TCM and AIL these ranged from ~0.28 to 0.41 and ~1.3 to 4.3, respectively. Overall, varying training sets and allocation strategies did not affect PA and MSE, with 52 non-overlapping and 0 overlapping genotypes per environment as the optimal cost-effective allocation framework. This suggests that implementing sparse testing designs could significantly reduce phenotyping costs by fivefold, without compromising PA in breeding programs for perennial crops such as Miscanthus.
高生物量多年生作物的表型鉴定工作繁重,且多年生作物育种项目的遗传增益速率通常较低。因此,发掘可提升育种流程效率的方法显得尤为关键。芒属植物(Miscanthus)是一种C4多年生草本植物,具备作为生物燃料及多种生物基产品原料的优良生物质生产特性。提升生物量产量可同时提高经济效益与环境效益,因此是芒属植物育种的核心目标之一。此外,要在多样环境条件下筛选出适应性优良的基因型,需开展多环境试验(multi-environment trials, METs)。稀疏测试(sparse testing)是一种基于基因组预测的策略,通过选取部分基因型在部分环境中进行表型鉴定,以此降低多环境试验的表型鉴定成本,随后可预测未观测的基因型-环境组合的性状表现。本研究针对由336份基因型组成的甜根子芒(Miscanthus sacchariflorus, MSA)群体展开分析,该群体在3种环境下进行了性状观测。本研究构建了3种预测模型,分别考虑主效应(环境、基因型、基因组)与互作效应(基因型-环境互作,即G×E互作),用于预测干生物量产量(dry biomass yield, YDY)、总茎秆数(total culm, TCM)、平均节间长度(average internode length, AIL)及茎秆节数(culm node number, CNN)。本研究设置了基于不同组成与规模的多套校正集,以在固定测试集规模的前提下,基于预测能力(predictive ability, PA)与均方误差(mean square error, MSE)评估模型性能。训练集规模介于52至112之间,用于预测3种环境下共224份未观测基因型的性状表现。研究结果显示,考虑G×E互作的模型在茎秆节数(CNN)与干生物量产量(YDY)上展现出最高的预测能力与最低的均方误差:CNN的PA约为0.77,MSE约为0.5;YDY的PA约为0.70,MSE约为1.3。而总茎秆数(TCM)与平均节间长度(AIL)的对应指标分别介于0.28至0.41、1.3至4.3之间。总体而言,不同的训练集设置与分配策略对预测能力与均方误差无显著影响,其中每个环境下选取52个非重叠、0个重叠基因型的分配方案为最优的成本效益框架。这表明,针对芒属植物这类多年生作物的育种项目,采用稀疏测试设计可在不降低预测能力的前提下,将表型鉴定成本降低五倍。



