强化学习
对抗制
计算机科学
机器学习
人工智能
水准点(测量)
过程(计算)
离线学习
样品(材料)
时间范围
缩小
元学习(计算机科学)
集成学习
外推法
在线和离线
光学(聚焦)
数据挖掘
集合预报
采样(信号处理)
国家(计算机科学)
数据建模
监督学习
地平线
作者
Hongye Cao,Fan Feng,Jing Huo,Shangdong Yang,Meng Fang,Tianpei Yang,Yang Gao
标识
DOI:10.1109/tnnls.2025.3636176
摘要
Model-based offline reinforcement learning (RL) constructs environment models from offline datasets to perform conservative policy optimization. Existing approaches focus on learning state transitions through ensemble models, rolling out conservative estimation to mitigate extrapolation errors. However, the static data makes it challenging to develop a robust policy, and offline agents cannot access the environment to gather new data. To address these challenges, we introduce Model-based Offline Reinforcement learning with AdversariaL data augmentation (MORAL). In MORAL, we replace the fixed horizon rollout by employing adversarial data augmentation to execute alternating sampling with ensemble models to enrich training data. Specifically, this adversarial process dynamically selects ensemble models against policy for biased sampling, mitigating the optimistic estimation of fixed models, thus robustly expanding the training data for policy optimization. Moreover, a differential factor (DF) is integrated into the adversarial process for regularization, ensuring error minimization in extrapolations. This data-augmented optimization adapts to diverse offline tasks without rollout horizon tuning, showing remarkable applicability. Extensive experiments on the D4RL benchmark demonstrate that MORAL outperforms other model-based offline RL methods in terms of policy learning and sample efficiency.
科研通智能强力驱动
Strongly Powered by AbleSci AI