遇见数据集

Learning When-to-Treat Policies

收藏
DataCite Commons2022-12-08 更新2024-07-28 收录
官方服务:

资源简介:

Many applied decision-making problems have a dynamic component: The policymaker needs not only to choose whom to treat, but also when to start which treatment. For example, a medical doctor may choose between postponing treatment (watchful waiting) and prescribing one of several available treatments during the many visits from a patient. We develop an “advantage doubly robust” estimator for learning such dynamic treatment rules using observational data under the assumption of sequential ignorability. We prove welfare regret bounds that generalize results for doubly robust learning in the single-step setting, and show promising empirical performance in several different contexts. Our approach is practical for policy optimization, and does not need any structural (e.g., Markovian) assumptions. Supplementary materials for this article are available online.

诸多实际决策问题均具有动态属性:决策者不仅需要选定干预对象,还需确定启动各项干预措施的时机。例如,在患者多次复诊过程中,医师可选择推迟干预(即观察等待),或是从数种可用干预方案中择一开具处方。我们提出一种「优势双重稳健(advantage doubly robust)」估计器,用于在「顺序可忽略性(sequential ignorability)」假设下,基于观测数据学习此类动态干预规则。我们证明了福利后悔界(welfare regret bounds),该结论推广了单阶段场景下双重稳健学习的既有成果,并在多种不同场景中展现出优异的实证性能。本方法可切实应用于政策优化场景,且无需任何结构性(如马尔可夫性(Markovian))假设。本文补充材料可在线获取。

提供机构:
Taylor & Francis
创建时间:
2020-10-06
二维码
社区交流群
二维码
科研交流群
商业服务