Benchmarking machine-learning models for emulating STICS grapevine yield across European wine regions
收藏资源简介:
Process-based crop models represent yield-forming mechanisms but become costly at ensemble scale. We tested whether interpretable machine learning could reproduce STICS simulations of healthy-crop grapevine yield constrained by climate, soil water and nitrogen while transferring to unseen locations and years. STICS simulations for Grenache, Cabernet Franc and Pinot Noir at 286 locations across European protected-designation-of-origin wine regions generated 15,260 site-years for 2004–2024. Models used 40 predictors selected within each fold from 144 climate, terrain and soil candidates, most available by 31 August, and were evaluated in 48 leave-location-and-time-out folds per cultivar. TabPFN was most accurate, with fold-median R² of 0.83 for Cabernet Franc, 0.82 for Pinot Noir and 0.51 for Grenache; nRMSE ranged from 7.0% to 11.3%. Grenache's lower R² reflected a larger relative error (nRMSE 11.3% versus 7.0–7.4%) rather than a narrower relative yield distribution, because its coefficient of variation matched that of the other cultivars. LightGBM remained within 0.03–0.06 R² of TabPFN and completed fitting and inference for three cultivar-specific models in 2.0 s, compared with an external 51-min serial STICS reference; its amortised full-batch inference was approximately 110,000 times faster. SHAP ranked extreme-heat frequency first for Grenache, whereas growing-degree-days and rooting depth dominated Cabernet Franc and Pinot Noir. STICS preserved broad observed regional yield rankings (Spearman ρ = 0.44, n = 52) but overestimated mean yield 2.31-fold and showed little interannual agreement. The emulators therefore enable rapid, interpretable analysis of STICS-simulated grapevine yield within the sampled environmental domain.



