竞争对手分析
计算机科学
质量(理念)
水准点(测量)
可用的
数据科学
业务
万维网
营销
地理
大地测量学
认识论
哲学
作者
Xiao Liu,Hao Yu,Hanchen Zhang,Yifan Xu,Xuanyu Lei,Hanyu Lai,Yu‐Cheng Gu,Hangliang Ding,Kaiwen Men,Kejuan Yang,Shudan Zhang,Xiang Deng,Aohan Zeng,Zhengxiao Du,Chenhui Zhang,Sheng Shen,Tianjun Zhang,Yu Su,Huan Sun,Minlie Huang
标识
DOI:10.48550/arxiv.2308.03688
摘要
The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively \textit{evaluate LLMs as agents} on challenging tasks in interactive environments. We present AgentBench, a multi-dimensional benchmark that consists of 8 distinct environments to assess LLM-as-Agent's reasoning and decision-making abilities. Our extensive test over \num API-based and open-sourced (OSS) LLMs shows that, while top commercial LLMs present a strong ability of acting as agents in complex environments, there is a significant disparity in performance between them and many OSS competitors that are no larger than 70B. We identify the typical reasons of failures in environments and LLMs, showing that poor long-term reasoning, decision-making, and instruction following abilities are the main obstacles for developing usable LLM agents. Improving instruction following and training on high quality multi-round alignment data could improve agent performance. And different from existing assumptions, training on code present ambivalent impacts on different agent tasks. Datasets, environments, and an integrated evaluation package for AgentBench are released at https://github.com/THUDM/AgentBench.
科研通智能强力驱动
Strongly Powered by AbleSci AI