已入深夜,您辛苦了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!祝你早点完成任务,早点休息,好梦!

Multi-armed Bandit Experimental Design: Online Decision-Making and Adaptive Inference

推论 计算机科学 多武装匪徒 人工智能 运筹学 机器学习 数学 后悔
作者
David Simchi‐Levi,Chonghuan Wang
出处
期刊:Management Science [Institute for Operations Research and the Management Sciences]
卷期号:71 (6): 4828-4846 被引量:3
标识
DOI:10.1287/mnsc.2023.00492
摘要

Multi-armed bandit has been well known for its efficiency in online decision-making in terms of minimizing the loss of the participants’ welfare during experiments (i.e., the regret). In clinical trials and many other scenarios, the statistical power of inferring the treatment effects (i.e., the gaps between the mean outcomes of different arms) is also crucial. Nevertheless, minimizing the regret entails harming the statistical power of estimating the treatment effect because the observations from some arms can be limited. In this paper, we investigate the trade-off between efficiency and statistical power by casting the multi-armed bandit experimental design into a minimax multi-objective optimization problem. We introduce the concept of Pareto optimality to mathematically characterize the situation in which neither the statistical power nor the efficiency can be improved without degrading the other. We derive a useful sufficient and necessary condition for the Pareto optimal solutions to the minimax multi-objective optimization problem. Additionally, we design an effective Pareto optimal multi-armed bandit experiment that can be tailored to different levels of the trade-off between the two objectives. Moreover, we extend the design and analysis to the setting where the outcome of each arm consists of an adversarial baseline reward and a stochastic treatment effect, demonstrating the robustness of our design. Finally, motivated by clinical trials, we examine the setting where the employed experiment must split the experimental units into a small number of batches, and we propose a flexible Pareto optimal design. This paper was accepted by George Shanthikumar, data science. Funding: The authors thank the Massachusetts Institute of Technology (MIT)-IBM partnership in Artificial Intelligence and the MIT Data Science Laboratory for support. Supplemental Material: The online appendix and data files are available at https://doi.org/10.1287/mnsc.2023.00492 .
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
xalone完成签到,获得积分10
刚刚
ccka关注了科研通微信公众号
刚刚
刚刚
1秒前
xalone发布了新的文献求助10
2秒前
动听酒窝关注了科研通微信公众号
2秒前
山间风发布了新的文献求助10
2秒前
榴莲嘎嘎发布了新的文献求助10
3秒前
qin完成签到,获得积分10
4秒前
深情安青应助pokexuejiao采纳,获得10
5秒前
7秒前
xxy发布了新的文献求助10
7秒前
侯康发布了新的文献求助10
8秒前
文静元霜发布了新的文献求助20
8秒前
9秒前
JAMES发布了新的文献求助30
10秒前
pony发布了新的文献求助10
10秒前
顾矜应助项目多多采纳,获得10
10秒前
11秒前
赘婿应助雷家采纳,获得10
12秒前
12秒前
12秒前
12秒前
汉堡包应助长夏采纳,获得10
14秒前
缓慢的柠檬关注了科研通微信公众号
15秒前
15秒前
ccka发布了新的文献求助10
16秒前
好运小陈发布了新的文献求助10
16秒前
顺利珂完成签到 ,获得积分10
17秒前
Dliii完成签到 ,获得积分10
17秒前
漱石枕流完成签到,获得积分10
17秒前
17秒前
脑洞疼应助xiao金采纳,获得10
18秒前
科研通AI6.4应助落后妖妖采纳,获得30
18秒前
英姑应助舒舒采纳,获得30
19秒前
FashionBoy应助开朗怜晴采纳,获得10
20秒前
研小白发布了新的文献求助10
20秒前
LiSiyi完成签到 ,获得积分10
21秒前
21秒前
22秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
化工安全与环保 1000
Autoparametric Resonance in Mechanical Systems 1000
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 800
Effects of Two Weeks of Red Light Therapy on Choroidal Thickness and Axial Length in Young Adults 700
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 600
Management and the Arts 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7656149
求助须知:如何正确求助?哪些是违规求助? 9226954
关于积分的说明 19827099
捐赠科研通 7222454
什么是DOI,文献DOI怎么找? 3280223
关于科研通互助平台的介绍 2440433
邀请新用户注册赠送积分活动 2279846