强化学习
计算机科学
贝叶斯定理
人工智能
受电弓
机器学习
钢筋
工程类
贝叶斯概率
结构工程
机械工程
作者
Hui Wang,Zhiwei Han,Xufan Wang,Yanbo Wu,Zhigang Liu
标识
DOI:10.1109/tte.2023.3293095
摘要
The fluctuation of pantograph–catenary contact force seriously affects the current collection quality, maintenance cost, and operation safety of high-speed trains. In recent years, agents developed with deep reinforcement learning (DRL) technology have achieved significant success. However, most of these jobs are restricted to narrow task distributions and stationary environments, which require a large amount of training data and cannot adapt quickly to the new task. We propose a contrastive learning-based Bayes-adaptive meta-reinforcement learning (CBAMRL) algorithm that addresses these limitations, enabling agents to learn new skills from a few transitions and adapt to the new environment. We first introduce a Bayes-adaptive training strategy, achieving zero-shot adaptation in nonstationary environments with high sample efficiency and competitive asymptotic performance. We proposed a contrastive learning-based contextual encoder to represent complex task distributions with similar structures, providing compact and sufficient task representation without modeling irrelevant dependencies. We evaluate the proposed method on a validated pantograph–catenary system (PCS) benchmark. Compared to the state-of-the-art DRL approach and traditional solutions, the experiment result demonstrates that the proposed algorithm can swiftly adapt to new operating circumstances and unknown perturbations with well-structured task representation and zero-shot adaptation.
科研通智能强力驱动
Strongly Powered by AbleSci AI