计算机科学
机器人
可执行文件
具身认知
稳健性(进化)
人工智能
标杆管理
感知
编码(社会科学)
人机交互
强化学习
内含代理
机器人学
概率逻辑
渲染(计算机图形)
仿人机器人
编码器
机器学习
可验证秘密共享
自主代理人
控制(管理)
可靠性(半导体)
视觉伺服
主动感知
冗余(工程)
先验概率
认知机器人学
演示式编程
瓶颈
作者
Max Fu,Justin Yu,Karim El-Refai,Ethan Kou,Haoru Xue,Huang Huang,Wenli Xiao,Guanzhi Wang,Fei-Fei Li,Guanya Shi,Jiajun Wu,Shankar Sastry,Yuke Zhu,Ken Goldberg,Linxi Fan
摘要
"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored. We present CaP-X, an open-access framework for systematically studying Code-as-Policy agents in robot manipulation. At its core is CaP-Gym, an interactive environment in which agents control robots by synthesizing and executing programs that compose perception and control primitives. Building on this foundation, CaP-Bench evaluates frontier language and vision-language models across varying levels of abstraction, interaction, and perceptual grounding. Across 12 models, CaP-Bench reveals a consistent trend: performance improves with human-crafted abstractions but degrades as these priors are removed, exposing a dependence on designer scaffolding. At the same time, we observe that this gap can be mitigated through scaling agentic test-time computation--through multi-turn interaction, structured execution feedback, visual differencing, automatic skill synthesis, and ensembled reasoning--substantially improves robustness even when agents operate over low-level primitives. These findings allow us to derive CaP-Agent0, a training-free framework that recovers human-level reliability on several manipulation tasks in simulation and on real embodiments. We further introduce CaP-RL, showing reinforcement learning with verifiable rewards improves success rates and transfers from sim2real with minimal gap. Together, CaP-X provides a principled, open-access platform for advancing embodied coding agents.
科研通智能强力驱动
Strongly Powered by AbleSci AI