强化学习
计算机科学
贝叶斯概率
机器学习
人工智能
分歧(语言学)
动态规划
数学优化
贝叶斯优化
唤醒睡眠算法
贪婪算法
算法
无监督学习
数学
泛化误差
语言学
哲学
作者
Fadil Santosa,Loren Anderson
标识
DOI:10.1109/icmla55696.2022.00106
摘要
We perform a comparison study on Bayesian sequential optimal experimental design algorithms applied to linear regression in two unknowns. We transform the Bayesian sequential optimal experimental design problem into a reinforcement learning problem to determine the power of deep reinforcement learning algorithms against baselines including batch design, greedy design, dynamic programming, and approximate dynamic programming. Using KL-divergence to measure information gain in the unknown parameters, we construct objectives for each algorithm to maximize information gain. This work showcases novel comparisons between the aforementioned algorithms and provides a new application of reinforcement learning to Bayesian sequential optimal experimental design for inverse problems in linear regression with multiple parameters.
科研通智能强力驱动
Strongly Powered by AbleSci AI