遇见数据集

Assessing the genotype-by-year effect on training set composition: scripts and dataset

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This dataset supports the study “Assessing the genotype-by-year effect on training set composition and alternatives to mitigate its impact on genomic selection accuracy.” The research investigates how genotype-by-year (G×Y) interactions influence genomic prediction (GP) accuracy and whether including overlapping checks and progenies in the training set can mitigate these effects. We hypothesized that limited overlap of entries across years would fail to correct for G×Y effects, reducing selection response. To test this, we conducted stochastic simulations using AlphaSimR, based on the structure of the LSU AgCenter rice breeding program. A total of 36 scenarios were simulated, combining four levels of G×Y interaction (0%, 25%, 50%, 75%) with three levels of overlap (0%, 5%, 10%) for both progenies and checks. The dataset includes: • Input parameters and configuration files for all scenarios • R scripts used for simulation and analysis • Output data containing genetic parameters: additive variance, population mean, best line performance, and prediction accuracy Results showed that stronger G×Y interactions consistently reduced GP accuracy and selection gains. Including up to 10% of checks and progenies in the training set did not significantly improve predictive performance. These findings indicate that such overlap is insufficient to account for temporal variation, especially in breeding programs with limited connectivity between cycles. The dataset enables full reproducibility and can support further research on training set optimization, modeling G×Y interactions, and evaluating selection strategies.

本数据集用于支撑题为《评估基因型-年份互作对训练集组成的影响及其缓解基因组选择精度负面影响的替代方案》的研究工作。该研究旨在探究基因型-年份(genotype-by-year, G×Y)互作如何影响基因组预测(Genomic Prediction, GP)精度,以及在训练集中纳入重叠参试材料与后代材料能否缓解此类互作带来的负面影响。 本研究提出假说:不同年份间参试材料的有限重叠无法校正G×Y效应,进而降低选择响应。为验证该假说,研究基于路易斯安那州立大学农业中心(LSU AgCenter)水稻育种项目的群体结构,利用AlphaSimR软件开展随机模拟实验。共设置36种模拟情景,将4个梯度的G×Y互作强度(0%、25%、50%、75%)与3个梯度的材料重叠率(0%、5%、10%)分别组合,其中重叠材料涵盖后代与参试对照两类。 本数据集包含以下内容: • 所有模拟情景的输入参数与配置文件 • 用于模拟与数据分析的R脚本 • 包含遗传参数的输出数据:加性方差、群体均值、最优品系表现及预测精度 研究结果显示,更强的G×Y互作会持续降低基因组预测精度与选择增益。在训练集中纳入最高10%的对照与后代材料,并未显著提升预测性能。上述结果表明,此类程度的材料重叠不足以抵消时间维度上的环境差异,尤其在各育种周期间连通性有限的育种项目中更是如此。 本数据集可实现研究的完全可重复性,并可为训练集优化、G×Y互作建模及选择策略评估等后续研究提供支撑。

创建时间:
2025-06-12
二维码
社区交流群
二维码
科研交流群
商业服务