一般化
计算机科学
生成语法
人工智能
生成模型
机器学习
任务(项目管理)
贝叶斯定理
机器人
钥匙(锁)
贝叶斯概率
数学
工程类
数学分析
计算机安全
系统工程
作者
Abhinav Agarwal,Sushant Veer,Allen Z. Ren,Anirudha Majumdar
标识
DOI:10.48550/arxiv.2111.08761
摘要
We are motivated by the problem of learning policies for robotic systems with rich sensory inputs (e.g., vision) in a manner that allows us to guarantee generalization to environments unseen during training. We provide a framework for providing such generalization guarantees by leveraging a finite dataset of real-world environments in combination with a (potentially inaccurate) generative model of environments. The key idea behind our approach is to utilize the generative model in order to implicitly specify a prior over policies. This prior is updated using the real-world dataset of environments by minimizing an upper bound on the expected cost across novel environments derived via Probably Approximately Correct (PAC)-Bayes generalization theory. We demonstrate our approach on two simulated systems with nonlinear/hybrid dynamics and rich sensing modalities: (i) quadrotor navigation with an onboard vision sensor, and (ii) grasping objects using a depth sensor. Comparisons with prior work demonstrate the ability of our approach to obtain stronger generalization guarantees by utilizing generative models. We also present hardware experiments for validating our bounds for the grasping task.
科研通智能强力驱动
Strongly Powered by AbleSci AI