Conformal Off-Policy Prediction in Contextual Bandits

DataCite Commons2026-01-07 更新2025-04-16 收录

下载链接：

https://service.tib.eu/ldmservice/dataset/905cb619-9a86-48da-a706-4438b968c29d

下载链接

链接失效反馈

官方服务：

资源简介：

Most off-policy evaluation methods for contextual bandits have focused on the expected outcome of a policy, which is estimated via methods that at best provide only asymptotic guarantees. However, in many applications, the expectation may not be the best measure of performance as it does not capture the variability of the outcome. In addition, particularly in safety-critical settings, stronger guarantees than asymptotic correctness may be required. To address these limitations, we consider a novel application of conformal prediction to contextual bandits. Given data collected under a behavioral policy, we propose conformal off-policy prediction (COPP), which can output reliable predictive intervals for the outcome under a new target policy.

提供机构：

TIB

创建时间：

2025-01-03

5,000+

优质数据集

54 个

任务类型

进入经典数据集