机器人
任务(项目管理)
计算机科学
人工智能
动作(物理)
杠杆(统计)
人机交互
平面图(考古学)
强化学习
自然语言理解
任务分析
机器人学
概化理论
自然语言
实用主义
认知机器人学
演示式编程
社交机器人
认知科学
机器人学习
自动计划和调度
对象(语法)
意外事件
iCub
行动计划
作者
Kaixian Qu,Guowei Lan,René Zurbrügg,Changan Chen,Christopher E. Mower,Haitham Bou-Ammar,Marco Hutter
标识
DOI:10.1109/lra.2026.3669804
摘要
Large language models (LLMs) have emerged as the dominant paradigm for robotic task planning using natural language instructions. However, trained on general internet data, LLMs are not inherently aligned with the embodiment, skill sets, and limitations of real-world robotic systems. Inspired by the emerging paradigm of verbal reinforcement learning—where LLM agents improve through self-reflection and few-shot learning without parameter updates—we introduce a framework that enables robots to learn task planning through real-world experience. mploys a vision-language model (VLM) as the robot's “brain” and “eye”, allowing it to visually evaluate action outcomes and self-reflect on failures. These reflections are stored in a short-term memory (STM), enabling the robot to quickly adapt its behavior during ongoing tasks. Upon task completion, the robot summarizes the lessons learned into its long-term memory (LTM). When facing new tasks, it can leverage retrieval-augmented generation (RAG) to plan more grounded action sequences by drawing on relevant past experiences and knowledge. Experiments on four challenging robotic tasks show that STM-based self-reflection increases task success rates from 35% to 84%, with emergent intelligent object interactions. In 12 real-world scenarios (including eight previously unseen tasks), the robot effectively learns from the LTM and improves single-trial success rates from 22% to 80%, with RAG outperforming naive prompting. These results highlight the effectiveness and generalizability of Project webpage: https://pragmabot.github.io/
科研通智能强力驱动
Strongly Powered by AbleSci AI