遇见数据集

MetaFlux: Meta-learning global carbon fluxes from sparse spatiotemporal observations

收藏
Zenodo2024-03-21 更新2026-05-26 收录
官方服务:

资源简介:

MetaFlux is a global, long-term carbon flux dataset of gross primary production and ecosystem respiration that is generated using meta-learning. The principle of meta-learning stems from the need to solve the problem of learning in the face of sparse data availability. Data sparsity is a prevalent challenge in climate and ecology science. For instance, in-situ observations tend to be spatially and temporally sparse. This issue can arise from sensor malfunctions, limited sensor locations, or non-ideal climate conditions such as persistent cloud cover. The lack of high-quality continuous data can make it difficult to understand many climate processes that are otherwise critical. The machine-learning community has attempted to tackle this problem by developing several learning approaches, including meta-learning that learns how to learn broad features across tasks to better infer other poorly sampled ones. In this work, we applied meta-learning to solve the problem of upscaling continuous carbon fluxes from sparse observations. Data scarcity in carbon flux applications is particularly problematic in the tropics and semi-arid regions, where only around 8–11% of long-term eddy covariance stations are currently operational. Unfortunately, these regions are important in modulating the global carbon cycle and its interannual variability. In general, we find that meta-trained machine models, including multi-layer perceptrons (MLP), long-short-term memory (LSTM), and bi-directional LSTM (BiLSTM), have lower validation errors on flux estimates by 9–16% when compared to their non-meta-trained counterparts. In addition, meta-trained models are more robust to extreme conditions, with 4–24% lower overall errors. Finally, we use an ensemble of meta-trained deep networks to generate a global product of ecosystem-scale photosynthesis and respiration fluxes from in-situ observations to daily and monthly global products at a 0.25-degree spatial resolution from 2001 to 2023, called "MetaFlux". We also checked for the seasonality, interannual variability, and correlation to solar-induced fluorescence of the upscaled product and found that MetaFlux outperformed state-of-the-art machine learning upscaling models, especially in critical semi-arid and tropical regions.

MetaFlux是一款全球性的长期碳通量数据集,涵盖总初级生产力(gross primary production)与生态系统呼吸(ecosystem respiration),其构建基于元学习(meta-learning)。元学习的原理源于解决数据稀缺场景下的学习难题,而数据稀疏性是气候与生态科学领域普遍存在的挑战。例如,原位观测(in-situ observations)往往在空间与时间维度上均存在稀疏性,该问题可能由传感器故障、布设位置有限,或是持续云量覆盖等非理想气候条件引发。高质量连续观测数据的缺失,会阻碍对诸多关键气候过程的认知。机器学习领域已开发多种学习方法以应对这一问题,其中元学习旨在学习跨任务的通用特征,从而更好地推断采样不足的任务。本研究将元学习应用于解决从稀疏观测中实现连续碳通量尺度扩展的问题。碳通量应用中的数据稀缺问题在热带与半干旱地区尤为突出,当前全球仅约8%~11%的长期涡度协方差(eddy covariance)观测站处于运行状态,而这些区域恰恰是调控全球碳循环及其年际变率的关键区域。研究结果表明,经元训练的机器学习模型,包括多层感知机(multi-layer perceptrons, MLP)、长短期记忆网络(long-short-term memory, LSTM)以及双向长短期记忆网络(bi-directional LSTM, BiLSTM),其通量估算的验证误差相较于非元训练的同类模型降低了9%~16%。此外,元训练模型对极端条件的鲁棒性更强,整体误差降低4%~24%。最后,本研究通过集成多个元训练深度网络,利用原位观测数据生成了生态系统尺度的光合与呼吸通量全球产品,该产品空间分辨率为0.25度,时间跨度为2001年至2023年,包含日尺度与月尺度数据,被命名为“MetaFlux”。我们还对该尺度扩展产品的季节特征、年际变率以及与日光诱导叶绿素荧光(solar-induced fluorescence)的相关性进行了验证,结果显示MetaFlux的表现优于现有最先进的机器学习尺度扩展模型,在关键的半干旱与热带地区优势尤为显著。

提供机构:
Zenodo
创建时间:
2024-03-21
二维码
社区交流群
二维码
科研交流群
商业服务