Hierarchical Low Rank Approximation of Likelihoods for Large Spatial Datasets
收藏资源简介:
Datasets in the fields of climate and environment are often very large and irregularly spaced. To model such datasets, the widely used Gaussian process models in spatial statistics face tremendous challenges due to the prohibitive computational burden. Various approximation methods have been introduced to reduce the computational cost. However, most of them rely on unrealistic assumptions for the underlying process and retaining statistical efficiency remains an issue. We develop a new approximation scheme for maximum likelihood estimation. We show how the composite likelihood method can be adapted to provide different types of hierarchical low rank approximations that are both computationally and statistically efficient. The improvement of the proposed method is explored theoretically; the performance is investigated by numerical and simulation studies; and the practicality is illustrated through applying our methods to two million measurements of soil moisture in the area of the Mississippi River basin, which facilitates a better understanding of the climate variability. Supplementary material for this article is available online.
气候与环境领域的数据集往往体量庞大且采样点分布不规则。针对此类数据集开展建模工作时,空间统计学中广泛应用的高斯过程(Gaussian process)模型会面临严峻挑战,其核心障碍在于计算负担难以承受。学界已提出多种近似方法以降低计算成本,但其中多数均对潜在过程(underlying process)做出了脱离实际的假设,且如何在简化计算的同时保留统计效率仍是尚未妥善解决的问题。为此,本文提出一种全新的最大似然估计(maximum likelihood estimation)近似框架。本文阐述了如何对复合似然(composite likelihood)方法进行适配,以构建兼具计算效率与统计效率的各类分层低秩近似方案。本文从理论层面分析了所提方法的性能优势,通过数值实验与仿真研究验证了方法的实际表现,并将所提方法应用于密西西比河流域的200万条土壤湿度观测数据,以此直观展示方法的实用价值,同时助力学界更深入地理解气候变率。本文的补充材料可在线获取。



