Multimodal fusion with vision-language-action models for robotic manipulation: A systematic review

计算机科学 人工智能 机器学习 标杆管理 稳健性(进化) 水准点(测量) 模块化设计 机器人学 融合机制 任务(项目管理) 软件部署 传感器融合 资源(消歧) 人机交互 融合 协议(科学) 在飞行中 计算模型 正确性 群机器人 分类学(生物学) 模态(人机交互)
作者
Muhayy Ud Din,Waseem Akram,Lyes Saad Saoud,Jan Rosell,Irfan Hussain
出处
期刊:Information Fusion [Elsevier BV]
卷期号:129: 104062-104062
标识
DOI:10.1016/j.inffus.2025.104062
摘要

• Provides a unified taxonomy that organizes more than 100 VLA architectures. • Maps 26 major VLA datasets using a framework based on task difficulty and modality richness. • Presents a large-scale quantitative analysis linking model design choices to normalized performance. • Demonstrates that diffusion-based decoders and hierarchical fusion significantly improve manipulation success. • Introduces the VLA-FEB benchmark with new metrics for measuring multimodal fusion quality and alignment. • Proposes an agentic VLA framework where LLM planners verify and re-plan actions using uncertainty-driven feedback for self-improving robotic autonomy. Vision Language Action (VLA) models represent a new frontier in robotics by unifying perception, reasoning, and control within a single multimodal learning framework. By integrating visual, linguistic, and action modalities, they enable multimodal fusion systems designed for instruction-driven manipulation and generalist autonomy. This systematic review synthesizes the state of the art in VLA research with an emphasis on architectures, algorithms, and applications relevant to robotic manipulation. We examine 102 models, 26 foundational datasets, and 12 simulation platforms, categorizing them according to their fusion strategies and integration mechanisms. Foundational datasets are evaluated using a novel criterion based on task complexity, modality richness, and dataset scale, allowing a comparative analysis of their suitability for generalist policy learning. We further introduce a structured taxonomy of fusion hierarchies and encoder-decoder families, together with a two-dimensional dataset characterization framework and a meta-analytic benchmarking protocol that quantitatively links design variables to empirical performance across benchmarks. Our analysis shows that hierarchical and late fusion architectures achieve the highest manipulation success and generalization, confirming the benefit of multi-level cross-modal integration. Diffusion-based decoders demonstrate superior cross-domain transfer and robustness compared to autoregressive heads. Dataset analysis highlights a persistent lack of benchmarks that combine high-complexity, multimodal, and long-horizon tasks, while existing simulators offer limited multimodal synchronization and real-to-sim consistency. To address these gaps, we propose the VLA Fusion Evaluation Benchmark to quantify fusion efficiency and alignment. Drawing on both academic and industrial advances, the review outlines future research directions in adaptive and modular fusion architectures, computational resource optimization, and the deployment of interpretable, resource-efficient robotic systems. We further propose a forward-looking agentic VLA paradigm where LLM planners integrate VLA skills as verifiable tools within a closed feedback loop for adaptive and self-improving robotic control. This work provides both a conceptual foundation and a quantitative roadmap for advancing embodied intelligence through multimodal information fusion across robotic domains. A public repository summarizing models, datasets, and simulators is available at: https://muhayyuddin.github.io/VLAs/ .

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
xxy完成签到 ,获得积分10
2秒前
水凝胶发布了新的文献求助10
2秒前
在水一方应助清秀斓采纳,获得50
3秒前
蓦然完成签到 ,获得积分10
3秒前
和谐凌波发布了新的文献求助10
3秒前
那小子真帅完成签到,获得积分10
5秒前
Miraitowa完成签到,获得积分20
6秒前
Owen应助火星上白羊采纳,获得10
6秒前
情怀应助Kar采纳,获得10
7秒前
zyzy完成签到,获得积分10
7秒前
8秒前
小二郎应助粗暴的君浩采纳,获得10
8秒前
hanwen完成签到,获得积分10
8秒前
molihuakai应助Zhao_JY采纳,获得10
9秒前
星辰大海应助小牛采纳,获得10
10秒前
11秒前
Ava应助MAJiaLu采纳,获得10
11秒前
烟花应助和谐凌波采纳,获得10
12秒前
斯文白白完成签到 ,获得积分20
12秒前
12秒前
12秒前
abbytcc完成签到,获得积分10
13秒前
啾比文发布了新的文献求助10
13秒前
14秒前
14秒前
14秒前
小蘑菇应助传统的半仙采纳,获得10
14秒前
共享精神应助mahuahua采纳,获得10
15秒前
CipherSage应助科研通管家采纳,获得10
16秒前
wangyy发布了新的文献求助10
16秒前
小徐完成签到,获得积分10
16秒前
今后应助科研通管家采纳,获得10
16秒前
16秒前
慕青应助科研通管家采纳,获得10
16秒前
16秒前
李健应助科研通管家采纳,获得10
17秒前
希成应助科研通管家采纳,获得10
17秒前
清秀斓发布了新的文献求助50
17秒前
隐形曼青应助科研通管家采纳,获得10
17秒前
搜集达人应助科研通管家采纳,获得100
17秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Reducing Compassion Fatigue, Secondary Traumatic Stress and Burnout 600
Comparative Elite Sport Development Systems, Structures and Public Policy 600
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Auslegungsgeschichte 500
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 500
What is the Future of Psychotherapy in Digital Age? Technology, AI Bots, and Psychotherapy after Covid 444
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7637373
求助须知:如何正确求助?哪些是违规求助? 9210973
关于积分的说明 19757588
捐赠科研通 7204676
什么是DOI,文献DOI怎么找? 3275647
关于科研通互助平台的介绍 2437328
邀请新用户注册赠送积分活动 2272834