Balanced Policy Evaluation and Learning

作者
Nathan Kallus
出处
期刊:Cornell University - arXiv [Cornell University]
被引量:19
标识
DOI:10.48550/arxiv.1705.07384
摘要

We present a new approach to the problems of evaluating and learning personalized decision policies from observational data of past contexts, decisions, and outcomes. Only the outcome of the enacted decision is available and the historical policy is unknown. These problems arise in personalized medicine using electronic health records and in internet advertising. Existing approaches use inverse propensity weighting (or, doubly robust versions) to make historical outcome (or, residual) data look like it were generated by a new policy being evaluated or learned. But this relies on a plug-in approach that rejects data points with a decision that disagrees with the new policy, leading to high variance estimates and ineffective learning. We propose a new, balance-based approach that too makes the data look like the new policy but does so directly by finding weights that optimize for balance between the weighted data and the target policy in the given, finite sample, which is equivalent to minimizing worst-case or posterior conditional mean square error. Our policy learner proceeds as a two-level optimization problem over policies and weights. We demonstrate that this approach markedly outperforms existing ones both in evaluation and learning, which is unsurprising given the wider support of balance-based weights. We establish extensive theoretical consistency guarantees and regret bounds that support this empirical success.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
山风完成签到,获得积分10
刚刚
SY1005完成签到 ,获得积分10
1秒前
1秒前
呀呀呀完成签到,获得积分10
2秒前
2秒前
星辰大海应助dm11采纳,获得10
3秒前
今后应助WangKai采纳,获得10
3秒前
4秒前
辛勤雨泽发布了新的文献求助10
4秒前
5秒前
顾矜应助阿萨德采纳,获得10
5秒前
赘婿应助Hannahhhhh采纳,获得10
6秒前
小二郎应助loria采纳,获得10
7秒前
听话的红牛完成签到,获得积分10
8秒前
8秒前
翟小暴发布了新的文献求助10
8秒前
呀呀呀发布了新的文献求助10
9秒前
烂漫悟空发布了新的文献求助10
9秒前
leefire关注了科研通微信公众号
9秒前
称心不尤发布了新的文献求助10
10秒前
旅行者完成签到,获得积分10
11秒前
情怀应助乐乐采纳,获得10
12秒前
DH发布了新的文献求助10
12秒前
13秒前
13秒前
研友_Lw4Ngn完成签到,获得积分10
13秒前
13秒前
spz150完成签到,获得积分10
14秒前
强砸完成签到,获得积分10
14秒前
小蘑菇应助糜灭龙采纳,获得30
15秒前
Qifan完成签到,获得积分10
15秒前
16秒前
16秒前
康康发布了新的文献求助10
16秒前
16秒前
17秒前
干净的琦应助cos采纳,获得10
18秒前
纵横无阙完成签到,获得积分10
19秒前
lzy发布了新的文献求助10
19秒前
高分求助中
Markov Chain Monte Carlo 10000
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Common Foundations of American and East Asian Modernisation: From Alexander Hamilton to Junichero Koizumi 5000
How to Use Machine Learning in Chemistry: An Introduction 1000
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Discerning Saints: Moralization of Intrinsic Motivation and Selective Prosociality at Work 500
Handbuch Trainingswissenschaft – Trainingslehre 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7583211
求助须知:如何正确求助?哪些是违规求助? 9161912
关于积分的说明 19605377
捐赠科研通 7165260
什么是DOI,文献DOI怎么找? 3266226
关于科研通互助平台的介绍 2431164
邀请新用户注册赠送积分活动 2257564