CLRD-GLPS: A Long-term Seasonal Dataset of Ruminant Livestock Distribution in China's Grazing Production Systems (2000-2021) Using Stacking-based Interpretable Machine Learning
收藏资源简介:
Understanding the spatial-temporal distribution of grazing livestock is crucial for assessing livestock system sustainability, managing animal diseases, mitigating climate change risks, and controlling greenhouse gas emissions. In China, grazing ruminants are predominantly distributed across vast grasslands in semi-humid and alpine regions. However, existing gridded livestock distribution datasets fail to distinguish between grazing and other livestock production systems and do not simultaneously account for long-term and seasonal dynamics. This study introduces CLRD-GLPS, a comprehensive dataset mapping China's ruminant livestock distribution in grazing livestock production systems from 2000 to 2021. Our approach addresses limitations in existing datasets by integrating interpretable machine learning methods to segment grazing livestock from total livestock populations and generate seasonal grazing pastures with dynamic grazing suitability masks. We developed a stacking-based ensemble methodology that enhances predictive performance while providing insights into distribution mechanisms. The stacking ensemble models demonstrate robust performance through 5-fold cross-validation, with R² values ranging from 0.909 to 0.967 for cattle and 0.874 to 0.914 for sheep. Validation results demonstrated the high accuracy of CLRD-GLPS across multiple spatial scales. At the county level, it strongly agreed with census data, effectively capturing grazing livestock distribution. City-level validation confirmed strong agreement (R² = 0.691–0.881), while grid-level validation using independent observations yielded R² = 0.79, further confirming the accuracy of fine-resolution predictions. The CLRD-GLPS dataset provides essential information for understanding grazing ruminant dynamics and developing targeted livestock management policies. Furthermore, our methodological framework offers a template for creating similar livestock distribution datasets for other regions and livestock production systems. This dataset is supported by the Second Tibetan Plateau Scientific Expedition and Research Program (STEP, grant no. 2019QZKK0906).
解析放牧家畜的时空分布格局,对评估家畜系统可持续性、防控动物疫病、减缓气候变化风险以及控制温室气体排放均具有重要意义。在中国,放牧反刍家畜主要分布于半湿润及高寒区域的广袤草原之中。然而,现有网格化家畜分布数据集无法区分放牧与其他家畜生产系统,且未同时兼顾长期与季节动态变化。本研究构建CLRD-GLPS数据集,这是一套覆盖2000至2021年、用于绘制中国放牧家畜生产系统内反刍家畜空间分布的综合数据集。本研究的方法通过集成可解释机器学习(Interpretable Machine Learning)方法,从总家畜种群中分割出放牧家畜,并生成带有动态放牧适宜性掩码(Mask)的季节放牧牧场,从而解决了现有数据集的局限性。我们开发了一种基于堆叠(Stacking)的集成学习方法,该方法在提升预测性能的同时,还能揭示家畜分布的形成机制。经5折交叉验证验证,该堆叠集成模型表现稳健:牛的决定系数(R²)介于0.909至0.967之间,绵羊的R²值则介于0.874至0.914之间。验证结果表明,CLRD-GLPS数据集在多空间尺度下均具备较高精度。在县域尺度上,该数据集与普查数据吻合度极高,能够有效捕捉放牧家畜的分布格局。市级尺度验证结果显示其吻合度同样优异(R²=0.691~0.881);而基于独立观测数据的格网尺度验证得到R²=0.79,进一步证实了该数据集高精度细分辨率预测的可靠性。CLRD-GLPS数据集为解析放牧反刍家畜的动态变化以及制定针对性的家畜管理政策提供了关键支撑数据。此外,本研究的方法框架可为其他区域及其他家畜生产系统构建同类家畜分布数据集提供参考范式。本数据集得到第二次青藏高原综合科学考察研究项目(STEP,项目编号:2019QZKK0906)的资助。



