水准点(测量)
计算机科学
强化学习
一套
人工智能
模仿
匹配(统计)
弹道
机器学习
机器人学
集合(抽象数据类型)
机器人
数学
程序设计语言
天文
考古
地理
物理
统计
历史
社会心理学
心理学
大地测量学
作者
Siddhant Haldar,Vaibhav Mathur,Denis Yarats,Lerrel Pinto
标识
DOI:10.48550/arxiv.2206.15469
摘要
Imitation learning holds tremendous promise in learning policies efficiently for complex decision making problems. Current state-of-the-art algorithms often use inverse reinforcement learning (IRL), where given a set of expert demonstrations, an agent alternatively infers a reward function and the associated optimal policy. However, such IRL approaches often require substantial online interactions for complex control problems. In this work, we present Regularized Optimal Transport (ROT), a new imitation learning algorithm that builds on recent advances in optimal transport based trajectory-matching. Our key technical insight is that adaptively combining trajectory-matching rewards with behavior cloning can significantly accelerate imitation even with only a few demonstrations. Our experiments on 20 visual control tasks across the DeepMind Control Suite, the OpenAI Robotics Suite, and the Meta-World Benchmark demonstrate an average of 7.8X faster imitation to reach 90% of expert performance compared to prior state-of-the-art methods. On real-world robotic manipulation, with just one demonstration and an hour of online training, ROT achieves an average success rate of 90.1% across 14 tasks.
科研通智能强力驱动
Strongly Powered by AbleSci AI