遇见数据集

Replication Data for: Inference at the Data's Edge: Gaussian Processes for Estimation and Inference in the Face of Extrapolation Uncertainty

收藏
DataONE2026-03-14 更新2026-05-19 收录
官方服务:

资源简介:

Many inferential tasks involve fitting models to observed data and predicting outcomes at new covariate values, requiring interpolation or extrapolation. Conventional methods select a single best-fitting model, discarding fits that were similarly plausible in-sample but would yield sharply different predictions out-of-sample. Gaussian Processes (GPs) offer a principled alternative. Rather than committing to one conditional expectation function, GPs deliver a posterior distribution over outcomes at any covariate value. This posterior effectively retains the range of models consistent with the data, widening uncertainty intervals where extrapolation magnifies divergence. In this way, the GP's uncertainty estimates reflect the implications of extrapolation on our predictions, helping to tame the \"dangers of extreme counterfactuals\" (King and Zeng 2006). The approach requires (i) specifying a covariance function linking outcome similarity to covariate similarity, and (ii) assuming Gaussian noise around the conditional expectation. We provide an accessible introduction to GPs with emphasis on this property, along with a simple, automated procedure for hyperparameter selection implemented in the R package gpss. We illustrate the value of GPs for capturing counterfactual uncertainty in three settings: (i) treatment effect estimation with poor overlap, (ii) interrupted time series requiring extrapolation beyond pre-intervention data, and (iii) regression discontinuity designs where estimates hinge on boundary behavior.

诸多推断任务均涉及基于观测数据拟合模型,并在新的协变量取值下预测结果,此过程需要插值或外推。传统方法仅选取单个最优拟合模型,舍弃那些在样本内拟合效果同样合理、但在样本外会产生显著不同预测结果的拟合模型。高斯过程(Gaussian Processes, GPs)提供了一种具备理论依据的替代方案。其无需绑定至单一条件期望函数,而是可为任意协变量取值下的结果生成后验分布。该后验分布可有效保留所有与数据相符的模型集合,并在因外推导致预测分歧放大的场景下拓宽不确定性区间。借此,高斯过程的不确定性估计可反映外推对预测结果的影响,有助于管控“极端反事实情形的风险”(King与Zeng,2006)。该方法需完成两项步骤:(i) 定义用于关联结果相似度与协变量相似度的协方差函数;(ii) 假设条件期望周围存在高斯噪声。本文对高斯过程进行了通俗易懂的介绍,重点阐述上述特性,并附带一套由R包gpss实现的简单自动化超参数选择流程。我们通过三类场景展示了高斯过程在捕捉反事实不确定性方面的应用价值:(i) 重叠性较差的处理效应估计;(ii) 需要在干预前数据范围外进行外推的间断时间序列分析;(iii) 估计结果依赖边界行为的断点回归设计。

创建时间:
2026-04-07
二维码
社区交流群
二维码
科研交流群
商业服务