遇见数据集

Raw dataset and R scripts for: Unravelling spatial drivers of topsoil total carbon variability in tropical paddy soils of Sri Lanka

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This data set represents the raw dataset, raster files associated with the environmental covariates used for modelling, and the R script that describes the flow of analyses used for the research article entitled: Unravelling the spatial drivers of topsoil total carbon concentration variability in paddy-growing soils in tropical agro-ecosystems of Sri Lanka. This study specifically aimed at identifying the spatial drivers and estimates of total carbon (TC) concentration in topsoil (0-15 cm) across the paddy-growing regions in tropical climates using Sri Lanka as a case study. Two distinct sampling strategies were used to collect soil samples for model calibration and validation purposes. For model calibration, a total of 888 sampling locations were sampled using a conditioned Latin Hypercube sampling approach. Additionally, 99 sampling sites were selected using a design-based stratified random strategy for independent evaluation of the developed models. Total carbon concentration (%) was analysed using an automated dry combustion method via a 2400 Series II CHN Elemental Analyser. Geospatial modelling of TC concentration was carried out through two distinct random forest models using a variety of environmental covariates. The environmental covariates used for the current analyses includes; mean annual rainfall (Rainfal_N), annual average mean temperature (Temp_N), annual average minimum temperature (Temp_Min_N), annual average maximum temperature (Temp_Max_N), vapour pressure deficient (VPD_N), MODIS enhanced vegetation index (Modis_N), SAGA wetness index (SAGA_WI_N), slope angle (Slope_d_N) and elevation (DEM_N). All environmental covariates were resampled to a spatial resolution of 100 m prior to spatial analysis. Furthermore, we deployed a novel area of applicability (AOA) calculation to quantify and identify regions where the current prediction is less reliable. In addition to AOA analysis, the uncertainty of TC prediction (%) was calculated at a 90% prediction interval. The influence of increasing the number of calibration sites on model prediction quality and reliability was assessed by using a user-defined sequence of calibration sites (e.g. n=200, n=300, n=400, n=500, n=600, n=700, n=800, n=888). For more information on the study area, sampling design, analytical data generation, modelling, and interpretation of the data, please refer to the Research article mentioned above.

本数据集包含原始数据集、建模所用环境协变量相关的栅格文件,以及为论文《解析斯里兰卡热带农业生态系统稻田土壤表层总碳浓度空间变异的驱动因子》所开发的分析流程R脚本。 本研究以斯里兰卡为案例区,旨在识别热带气候下稻田区域表层土壤(0~15 cm)总碳(TC)浓度的空间驱动因子及其估算方法。 本研究采用两种不同的采样策略采集土壤样本,用于模型校准与验证。模型校准阶段,采用条件拉丁超立方采样(conditioned Latin Hypercube sampling)方法共布设888个采样点。此外,采用基于设计的分层随机策略选取99个采样点,用于所构建模型的独立验证。 本研究采用2400 Series II型CHN元素分析仪的自动干烧法测定总碳浓度(%)。基于多种环境协变量,采用两种不同的随机森林(random forest)模型开展总碳浓度的空间地理建模。本次分析所使用的环境协变量包括:年平均降雨量(Rainfal_N)、年平均气温(Temp_N)、年平均最低气温(Temp_Min_N)、年平均最高气温(Temp_Max_N)、水汽压亏缺(VPD_N)、MODIS增强植被指数(Modis_N)、SAGA湿度指数(SAGA_WI_N)、坡度角(Slope_d_N)以及数字高程模型(DEM_N)。所有环境协变量在空间分析前均重采样至100米的空间分辨率。 此外,本研究采用一种新型适用区域(Area of Applicability, AOA)计算方法,对当前预测可靠性较低的区域进行量化与识别。除适用区域分析外,本研究还以90%预测区间计算总碳浓度预测结果的不确定性(%)。通过设置用户自定义的校准样点数量序列(如n=200、n=300、n=400、n=500、n=600、n=700、n=800、n=888),评估校准样点数量增加对模型预测质量与可靠性的影响。如需了解研究区、采样设计、分析数据生成、建模及数据解读的更多细节,请参阅上述研究论文。

创建时间:
2023-12-20
二维码
社区交流群
二维码
科研交流群
商业服务