Dataset for: Multi-Data Source-Based Machine Learning Modelling Framework for Remote Estimation of Soil Organic Carbon and Carbon Credits Validation
收藏资源简介:
Dataset used for the scientific work entiled "Multi-Data Source-Based Machine Learning Modelling Framework for Remote Estimation of Soil Organic Carbon and Carbon Credits Validation" The estimation of the soil organic carbon (SOC) using remotely sensed data is playing an increasingly important role in agro-environmental studies, particularly as climate change exerts growing pressure on all anthropogenic activities. Testing and validating a system capable of implementing multi-data sources with artificial intelligence algorithms to predict SOC in space and time is crucial, as SOC controls and regulates various physical, chemical, and biological processes in agriculture. Many studies focus on developing SOC models that incorporate numerous covariates. However, for reasons of cost, time effectiveness and scalability, it would be preferable to utilize only a few variables selected ad hoc based on their importance. The objective of this study is to evaluate the performance of three algorithms (Linear, XGBoost, and Random Forest) in predicting SOC across three sites characterized by diverse crops, soil types, and irrigation systems, employing machine learning techniques. Two distinct datasets were employed to attain this objective. Both datasets encompass information on i) crop types; ii) Normalized Difference Vegetation Index maps derived from five satellites (Sentinel-2, Copernicus satellite system) during the crop specific maximum peak of Leaf Area Index; iii) climatic data; iv) soil type defined according to the World Reference Base classification; and v) irrigation system. In one of the datasets, an additional artificial covariate, namely the ‘zone management’ categorical variable, was introduced. This variable was automatically generated using satellite images and cluster analysis for each study site, enhancing the dataset with a unique artificial feature. Regardless of the use of the Linear, XGBoost, or Random Forest algorithm, the utilization of the dataset containing the artificial covariate, ‘zone management,’ exhibited the best performance, resulting in an average accuracy improvement of 40% compared to the dataset lacking this covariate for the tested models. For future SOC modelling studies, it is advisable to explore the potential of incorporating artificial zone management covariate coupled with selected agronomic variables to enhance predictive capabilities.



