摘要
Embodied intelligence, as an interdisciplinary frontier that integrates artificial intelligence and robotics, relies fundamentally on robust task planning capabilities to enable autonomous decision-making in dynamic and complex environments. This paper systematically reviews the methods of embodied intelligence task planning that are grounded in large language models (LLMs), conducting a comprehensive multidimensional analysis across theoretical frameworks, technical pathways, and multi-agent collaboration dynamics. The study reveals that LLMs, through their advanced mechanisms of natural language comprehension, multimodal fusion, and dynamic reasoning, effectively address the inherent limitations of conventional rule-based systems in terms of flexibility, adaptability, and interactivity. The array of single-agent strategies, which encompasses end-to-end planning, phased planning, and dynamic planning, along with the diverse multi-agent collaborative frameworks involving centralized, distributed, and hybrid architectures, collectively establish novel paradigms for task decomposition and execution in intricate real-world scenarios. Furthermore, the paper critically examines key challenges like logical consistency, energy efficiency, and safety verification that impede the widespread use of LLM-based systems, while proposing future directions such as neuro-symbolic hybrid architectures and lightweight model optimization to boost practicality. Notably, these advancements not only strengthen the technical foundations of embodied intelligence but also facilitate smoother integration of intelligent agents into daily life and industrial environments.