Data and Code for: Image-based deep learning enables flexible, holistic 3D mapping of multiple soil properties
收藏资源简介:
This data release accompanies a manuscript comparing deep learning architectures for predicting soil properties across the Upper Colorado River Basin, USA. The project evaluates convolutional neural network, Vision Transformer Autoencoder, and Masked Autoencoder architectures against a random forest baseline for predicting 8 soil properties (sand, silt, clay, CaCO3, EC, SAR, pH, SOC) at 7 standard depth intervals (0, 5, 15, 30, 60, 100, and 150 cm), totaling 56 prediction targets. This is the 3nd version of the release for a minor revision that updates files in the scripts.zip and outputs.zip that follow the same structure as version 2. The release contains all source code (R and Python scripts), model outputs (metrics, predictions, training histories), spatial prediction maps (GeoTIFF), environmental covariate raster layers, and ancillary vector datasets used in the study. All deep learning models were developed using TensorFlow/Keras in R, with Python scripts used for HPC-based pretraining on the USDA SCINet Atlas cluster and for SHAP variable importance analysis. Due to privacy agreements, only 1854 out of 2042 soil observations used in the study could be included in the data release here. Please see the Data_release_README.docx file for complete explanation of all available files. This paper has been revised and re-submitted to the journal Geoderma, and the status will be updated as it goes through review.



