遇见数据集

Data for variation of magnesium drives plant adaption to heterogeneous environments by regulating efficiency in photosynthesis on a large scale

收藏
NIAID Data Ecosystem2026-05-02 收录
官方服务:

资源简介:

Magnesium (Mg) is a vital nutrient for plants, and its role in photosynthesis, enzyme regulation, and resistance to environmental stress is becoming increasingly evident. However, there is a paucity of knowledge regarding the characteristics of Mg (content, density, and stock) on a large scale, particularly at the community level, which serves as a fundamental unit for linking to ecosystem functions. A leaf-branch-trunk-root-matched database of the Mg content (mg g–1) and biomass (g m–2) of plant organs across 1972 sampling sites in China was constructed based on field surveys and data compilation. Using machine learning algorithms, we comprehensively explored the spatial patterns and main influencing factors of plant Mg content and density (g m–2). Deserts exhibited higher Mg content, with the primary influencing factors being high temperature and soil Mg supply. High Mg density values occurred in forests. The spatial patterns of Mg content and density underscore the adaptation of plants to environmental changes and nutrient retention capacity of forests, respectively. Our research not only provides valuable information on the distribution of Mg across different communities. Methods When investigating community structure, we set up 3 or 4 sample plots for trees (20 m × 20 m), 6 for shrubs (5 m × 5 m), and 8 for herbs (1 m × 1 m). The height and breast height diameter for trees, height, basal diameter, and crown width for shrubs were investigated for biomass calculation using allometric growth equations. We collected the aboveground portion of herbs by species after drying and weighing. The biomass of trees and shrubs was calculated, and the aboveground biomass of herbs was obtained by weighing, whereas the belowground biomass was calculated using the root-crown ratio. Healthy and fully expanded leaves and top branches (diameter < 1 cm) were collected for trees and shrubs, along with core samples at breast height (trees only). All aboveground portions of the herbs were collected. Intact fine roots (diameter < 2 mm) were excavated along lateral roots. Based on the results of the community structure investigation, samples from different species were proportionally mixed to determine the Mg content of the plant communities. Soil samples of depth of 0–10 cm were collected in each plot to determine soil properties. After thorough washing with distilled water, plant samples were placed in a drying oven at 60°C until a constant mass was reached. For soil samples, plant roots and gravel were first removed, sieved through a 2 mm mesh, and naturally air-dried. Subsequently, all plant and soil samples were ground into a fine powder using an agate mortar (RM200, Retsch, Haan, Germany) and a ball mill (RM200, Retsch). The powdered samples were digested in a microwave system (Mars X press, CEM, Matthews, NC, USA) and then analyzed for Mg content (mg g−1) using an inductively coupled plasma spectrometer (ICP-OES, Optima 5300 DV, Perkin Elmer, Waltham, MA, USA). Ten environmental factors were chosen for machine learning and then predicting China Mg content and density, including five climate factors (MAT, TWMax, TCMin, MAP, and AI), two vegetation factors (PAR and NDVI), and three soil factors (SpH, CEC, and ExcMg). Data for MAT, maximum temperature of the warmest month (TWMax, °C), minimum temperature of the coldest month (TCMin, °C), and MAP were obtained from the WorldClim database (https://worldclim.org/). The aridity index (AI) was sourced from the CGIAR-CSI database (https://csidotinfo.wordpress.com/). Photosynthetically active radiation (PAR, W m–2) from 2000 to 2018 was collected from the National Tibetan Plateau Data Center (https://data.tpdc.ac.cn/) and interpolated to 1km using the Kriging method. Data on the normalized difference vegetation index (NDVI) from 2000 to 2018 were extracted from the Resource and Environment Science and Data Center (https://www.resdc.cn/). Soil pH (SpH), cation exchange capacity (CEC, cmol kg–1), and soil exchangeable Mg content (ExcMg, me 100g–1) were extracted from the Big Earth Data for Three Poles (https://poles.tpdc.ac.cn/). Maximum rate of Rubisco carboxylation (VCmax, μmol m–2 s–1) was extracted from the National Ecosystem Science Data Center (https://nesdc.org.cn/) to verify the relationship between leaf photosynthetic capacity and Mg content. In addition, vegetation types (broadleaf forest, needleleaf, and broadleaf mixed forest, needleleaf forest, shrub, meadow, steppe, tussock, and desert) extracted from the vegetation map of China was used to predict Mg content and density. RF, boosted regression trees (BRT), extreme gradient boosting (XGBoost), and light gradient boosting machine (LightGBM) algorithms were used to develop models that predict the Mg content and density in the leaves, branches, trunks, and roots, as well as the aboveground, belowground, and total aspects of vegetation across China. Before training the models for Mg density, the data were split into a training set (70%) and validation set (30%). For each model, the optimal combination of hyperparameters was identified through a grid search and 5-fold cross-validation, and the predictive performance of the model under the optimal hyperparameters was then evaluated on the validation set. The evaluation metrics included the root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R2). A lower RMSE and MAE and higher R2 indicate a better predictive performance of the model. To achieve more robust prediction results and assess the uncertainty of machine learning model predictions, we used bootstrapped datasets for training the models. For samples larger than 100, we generated datasets of the same size as the original dataset; for samples smaller than 100, we generated datasets with 100 samples. We performed bootstrapping to generate 100 new datasets for each dataset. Machine learning models were trained and tuned on these new datasets. Performance was assessed as the average values of 100 models. The uncertainty of the predictions was evaluated using the coefficient of variation, and the final prediction was the average of 100 models. For each dataset, the algorithm with the lowest RMSE was selected for prediction. RF consistently achieved the lowest RMSE across all datasets and was therefore chosen for the final predictions.

镁(Magnesium, Mg)是植物必需的关键营养元素,其在光合作用、酶调控以及环境胁迫抗性中的作用日益受到关注。然而,当前针对大尺度下植物镁的特征(含量、密度与储量),尤其是作为联结生态系统功能基本单元的群落水平的相关研究仍较为匮乏。本研究基于野外调查与数据汇编,构建了覆盖中国1972个采样点的植物器官镁含量(mg g–1)与生物量(g m–2)的叶-枝-干-根匹配数据库。通过机器学习算法,我们全面探究了植物镁含量与密度(g m–2)的空间格局及其主要影响因子。研究发现荒漠区域的植物镁含量更高,其主要影响因子为高温与土壤镁供给;森林区域的镁密度值较高。植物镁含量与密度的空间格局分别反映了植物对环境变化的适应策略与森林的养分保留能力。本研究为不同群落类型的镁分布提供了宝贵的基础数据。 方法 在调查群落结构时,我们针对乔木设置3或4个20 m × 20 m的样方,灌木设置6个5 m × 5 m的样方,草本设置8个1 m × 1 m的样方。为计算生物量,我们测定了乔木的树高与胸径,灌木的树高、基径与冠幅,并采用异速生长方程进行估算。针对草本植物,我们按物种采集其地上部分,经烘干称重后获取生物量。最终分别计算得到乔木与灌木的生物量,通过称重获取草本地上生物量,而地下生物量则通过根冠比进行估算。 针对乔木和灌木,采集健康且完全展开的叶片以及直径<1 cm的顶生枝条,同时采集乔木的胸径处木芯样本;采集草本所有地上部分;沿侧根挖掘完整的细根(直径<2 mm)。基于群落结构调查结果,按比例混合不同物种的样品,以测定植物群落的镁含量。每个样方采集0–10 cm深度的土壤样品以测定土壤理化性质。 植物样品经蒸馏水彻底清洗后,置于60℃烘箱中烘干至恒重。土壤样品则先去除根系、砾石,过2 mm筛后自然风干。随后,使用玛瑙研钵(RM200,莱驰,德国哈恩)和球磨仪(RM200,莱驰)将所有植物和土壤样品研磨为细粉。将粉末样品置于微波消解系统(Mars X press,CEM,美国北卡罗来纳州马修斯)中消解,随后使用电感耦合等离子体发射光谱仪(Inductively Coupled Plasma Optical Emission Spectrometer, ICP-OES,Optima 5300 DV,珀金埃尔默,美国马萨诸塞州沃尔瑟姆)测定镁含量(mg g−1)。 选取10项环境因子用于机器学习建模以预测中国区域的植物镁含量与密度,其中包括5项气候因子(年均温MAT、最热月最高温TWMax、最冷月最低温TCMin、年降水量MAP、干燥度指数AI),2项植被因子(光合有效辐射PAR、归一化植被指数NDVI),以及3项土壤因子(土壤pH SpH、阳离子交换量CEC、交换性镁含量ExcMg)。年均温、最热月最高温、最冷月最低温及年降水量数据取自WorldClim数据库(https://worldclim.org/)。干燥度指数AI源自CGIAR-CSI数据库(https://csidotinfo.wordpress.com/)。2000–2018年的光合有效辐射数据取自国家青藏高原科学数据中心(https://data.tpdc.ac.cn/),并通过克里金法插值至1 km分辨率。2000–2018年的归一化植被指数数据取自资源与环境科学与数据中心(https://www.resdc.cn/)。土壤pH、阳离子交换量及土壤交换性镁含量数据取自三极大数据科学中心(https://poles.tpdc.ac.cn/)。核酮糖-1,5-二磷酸羧化酶/加氧酶(Rubisco)最大羧化速率(VCmax,μmol m–2 s–1)取自国家生态系统科学数据中心(https://nesdc.org.cn/),用于验证叶片光合能力与镁含量的关联关系。此外,取自中国植被图的植被类型数据(包括阔叶林、针叶林、针阔混交林、针叶林、灌丛、草甸、草原、草丛与荒漠)被用于预测植物镁含量与密度。 采用随机森林(Random Forest, RF)、提升回归树(Boosted Regression Trees, BRT)、极端梯度提升(Extreme Gradient Boosting, XGBoost)及轻量级梯度提升机(Light Gradient Boosting Machine, LightGBM)算法,构建模型以预测中国区域植物各器官(叶、枝、干、根)以及植被地上、地下及总部分的镁含量与密度。在训练镁密度预测模型前,将数据划分为训练集(70%)与验证集(30%)。针对每个模型,通过网格搜索与5折交叉验证确定最优超参数组合,随后在验证集上评估最优超参数下模型的预测性能。评估指标包括均方根误差(RMSE)、平均绝对误差(MAE)与决定系数(R²):RMSE与MAE越低、R²越高,代表模型预测性能越好。为获得更稳健的预测结果并评估机器学习模型预测的不确定性,我们使用自助采样数据集训练模型:对于样本量大于100的数据集,生成与原数据集规模一致的自助数据集;对于样本量小于100的数据集,生成规模为100的自助数据集。针对每个原始数据集,我们通过自助采样生成100个新数据集,在这些新数据集上训练并调优机器学习模型。模型性能以100次模型训练结果的平均值进行评估,预测不确定性通过变异系数进行评估,最终预测结果为100次模型预测的平均值。针对每个数据集,选取RMSE最低的算法进行预测。随机森林在所有数据集上均取得最低的RMSE,因此被选为最终预测的算法。

创建时间:
2024-08-22
二维码
社区交流群
二维码
科研交流群
商业服务