Semantically Guided Task Planning: Supervised Vision-Language-Action Model by Large Language Models

计算机科学 可执行文件 机器人 任务(项目管理) 人工智能 编码器 任务分析 机器学习 图形 生成模型 动作(物理) 人机交互 语言模型 编码(集合论) 自然语言处理 建模语言 语义学(计算机科学) 可视化 语言理解 生成语法 自然语言 语义映射 形势意识 数据建模 上下文模型
作者
X L Li,Guohui Tian,Yongcheng Cui
出处
期刊:IEEE Transactions on Circuits and Systems for Video Technology [Institute of Electrical and Electronics Engineers]
卷期号:36 (5): 7411-7425
标识
DOI:10.1109/tcsvt.2025.3642702
摘要

Enabling robots to perform everyday tasks has become increasingly important. Task planning, which decomposes task instructions into executable action sequences, is crucial for equipping robots with the ability to handle daily activities. Currently, there are two main effective methods for task planning: one relies on the reasoning capabilities of Large Language Models (LLMs), but it struggles with handling the underlying motion. The other is based on the generative capabilities of Vision-Language-Action (VLA) model, which often lacks essential semantic details. To overcome these limitations, this paper introduces a novel Semantically Supervised Vision-Language-Action (SS-VLA) model. This model addresses the constraints of previous method that relied solely on single-frame image by designing an adaptive visual sequence encoder that integrates continuous visual streams. This encoder efficiently captures and integrates multi-scale spatial and temporal features from the robot’s first-person visual perspective. Furthermore, the model utilizes LLMs to decompose task instructions into subtasks and organize them into graph structure, using Graph Attention Network (GAT) to extract features from subtask sequences and supervise the generation of action sequences. This method not only enhances the alignment of actions with task instructions but also ensures the contextual and semantic accuracy of the robot’s activities, significantly enhancing the task execution capabilities of robots in complex environments. We evaluated our model on the ALFRED and TEACh benchmark, achieving higher performance compared to existing methods, especially in unseen scenes. Additionally, we successfully deployed our model in the AI2-THOR virtual environment and on the TIAGo real robot, demonstrating the effectiveness of our method. Our code is available at: https://github.com/Li-XD-Pro/SS-VLA.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
江南第八完成签到,获得积分10
1秒前
123完成签到,获得积分10
1秒前
1秒前
言目木完成签到,获得积分10
1秒前
儒雅的若翠完成签到,获得积分0
1秒前
玄音完成签到,获得积分10
1秒前
科研通AI6.4应助沫茉采纳,获得10
1秒前
毕院士完成签到,获得积分10
1秒前
llk完成签到,获得积分10
1秒前
Violeta完成签到,获得积分20
1秒前
在水一方应助LZZ采纳,获得10
1秒前
shenxixi完成签到,获得积分10
1秒前
liangxianli完成签到,获得积分10
1秒前
小王完成签到,获得积分10
1秒前
qipao完成签到,获得积分10
2秒前
快乐的烨磊完成签到,获得积分10
2秒前
海蓝云天完成签到,获得积分0
2秒前
斯文败类应助海绵宝宝采纳,获得10
2秒前
lixj0419完成签到,获得积分10
2秒前
炸薯条大王完成签到,获得积分10
3秒前
Wells完成签到,获得积分10
3秒前
yz完成签到,获得积分10
3秒前
无极微光应助科研通管家采纳,获得20
3秒前
清脆山槐完成签到,获得积分10
3秒前
liu完成签到,获得积分10
3秒前
3秒前
长情的雨完成签到,获得积分10
3秒前
3秒前
day完成签到,获得积分10
3秒前
慕青应助科研通管家采纳,获得10
3秒前
科研狗应助科研通管家采纳,获得50
3秒前
科研通AI2S应助科研通管家采纳,获得10
4秒前
4秒前
科研狗应助科研通管家采纳,获得50
4秒前
dawd12完成签到,获得积分10
4秒前
Applejuice完成签到,获得积分10
4秒前
Akim应助科研通管家采纳,获得10
4秒前
Orange应助科研通管家采纳,获得20
4秒前
liyi完成签到,获得积分10
4秒前
充电宝应助科研通管家采纳,获得10
4秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
DIPPR Project 801 - Full Version 380
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7766181
求助须知:如何正确求助?哪些是违规求助? 9310092
关于积分的说明 20315074
捐赠科研通 7351008
什么是DOI,文献DOI怎么找? 3315033
关于科研通互助平台的介绍 2464576
邀请新用户注册赠送积分活动 2329603