强化学习
钢筋
计算机科学
自然语言
自然(考古学)
认知心理学
人工智能
自然语言处理
认知科学
心理学
人机交互
社会心理学
生物
古生物学
作者
Shinya Masadome,Taku Harada
摘要
Reinforcement learning (RL) has found applications across diverse domains; however, it grapples with challenges when formulating reward functions and exhibits low exploration efficiency. Recent studies leveraging large language models (LLMs) have made strides in addressing these issues. However, for RL agents to be practically deployable, elucidating their decision‐making process is crucial for enhancing explainability. We introduce a novel RL approach aimed at alleviating the burden of designing reward functions and facilitating natural language explanations for actions grounded in the agent's decisions. Our method employs two types of agents: a low‐level agent responsible for concrete action selection and a high‐level agent tasked with setting abstract action goals. The high‐level agent undergoes training using a hybrid reward function framework, which incentivizes its actions by comparing them with those generated by an LLM across discretized states. Meanwhile, the training of the low‐level agent is guided by a reward function designed using the EUREKA algorithm. We applied the proposed method to the cart‐pole problem and demonstrated its ability to achieve a learning convergence rate while reducing human effort. Moreover, our approach yields coherent natural language explanations elucidating the rationale behind the agent's actions. © 2025 The Author(s). IEEJ Transactions on Electrical and Electronic Engineering published by Institute of Electrical Engineers of Japan and Wiley Periodicals LLC.
科研通智能强力驱动
Strongly Powered by AbleSci AI