仿人机器人
机器人
计算机科学
人工智能
人机交互
模仿
基础(证据)
机器人学
机器人控制
方案(数学)
机器人运动学
控制工程
变压器
iCub
移动机器人
群机器人
模拟
认知机器人学
作者
Nvidia Nvidia,:,Johan Björck,Fernando Castañeda,Nikita Cherniadev,Xingye Da,Runyu Ding,Linxi Fan,Fang Yu,Dieter Fox,Fengyuan Hu,Shanshan Huang,Joel Jang,Zhenyu Jiang,Jan Kautz,Kaushil Kundalia,Lixing Lao,Z. Y. Li,Zongyu Lin,Kevin Lin
标识
DOI:10.48550/arxiv.2503.14734
摘要
General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for enabling the robots to reason about novel situations, robustly handle real-world variability, and rapidly learn new tasks. To this end, we introduce GR00T N1, an open foundation model for humanoid robots. GR00T N1 is a Vision-Language-Action (VLA) model with a dual-system architecture. The vision-language module (System 2) interprets the environment through vision and language instructions. The subsequent diffusion transformer module (System 1) generates fluid motor actions in real time. Both modules are tightly coupled and jointly trained end-to-end. We train GR00T N1 with a heterogeneous mixture of real-robot trajectories, human videos, and synthetically generated datasets. We show that our generalist robot model GR00T N1 outperforms the state-of-the-art imitation learning baselines on standard simulation benchmarks across multiple robot embodiments. Furthermore, we deploy our model on the Fourier GR-1 humanoid robot for language-conditioned bimanual manipulation tasks, achieving strong performance with high data efficiency.
科研通智能强力驱动
Strongly Powered by AbleSci AI