A Survey on Evaluation of Large Language Models

人气 工程伦理学 心理学 工程类 社会心理学
作者
Yupeng Chang,Xu Wang,Jindong Wang,Yuan-Hsuan Wu,Linyi Yang,Kaijie Zhu,Hao Chen,Xiaoyuan Yi,Cunxiang Wang,Yidong Wang,Wei Ye,Yue Zhang,Yi Chang,Philip S. Yu,Qiang Yang,Xing Xie
出处
期刊:Cornell University - arXiv [Cornell University]
被引量:200
标识
DOI:10.48550/arxiv.2307.03109
摘要

Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate, where to evaluate, and how to evaluate. Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, educations, natural and social sciences, agent applications, and other areas. Secondly, we answer the `where' and `how' questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
miao发布了新的文献求助10
刚刚
cc完成签到 ,获得积分10
1秒前
小馒头发布了新的文献求助10
1秒前
2秒前
我爱化学完成签到 ,获得积分10
2秒前
鹿璟璟完成签到 ,获得积分10
3秒前
麦苗果果发布了新的文献求助50
3秒前
3秒前
3秒前
5秒前
5秒前
MAXDONE完成签到,获得积分10
6秒前
6秒前
Hello应助BrogirlMiku采纳,获得10
6秒前
CFF发布了新的文献求助10
6秒前
7秒前
7秒前
苏某坡完成签到,获得积分10
8秒前
科研通AI6.4应助毛学腾采纳,获得10
8秒前
小葡萄完成签到 ,获得积分10
8秒前
wg发布了新的文献求助10
9秒前
上善若水完成签到 ,获得积分10
9秒前
10秒前
受伤蝴蝶发布了新的文献求助10
10秒前
茹茹发布了新的文献求助10
10秒前
紫炫完成签到 ,获得积分10
11秒前
xulaoshi发布了新的文献求助10
11秒前
小马甲应助张宇鑫采纳,获得10
12秒前
13秒前
华仔应助唐科研采纳,获得10
13秒前
小巧紫易完成签到,获得积分20
13秒前
13秒前
crystal完成签到 ,获得积分10
13秒前
充电宝应助菠萝冰采纳,获得10
13秒前
人才发布了新的文献求助10
15秒前
15秒前
15秒前
16秒前
17秒前
小巧紫易发布了新的文献求助10
17秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Geist der Kunst und Kultur 1000
Resistance Spot Welding Dataset for Automobile Body-in-White Quality Analysis 748
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 700
Child and Adolescent Psychology 600
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
丝光沸石活性位点定向调控及其二甲醚羰基化性能研究 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7418606
求助须知:如何正确求助?哪些是违规求助? 9022377
关于积分的说明 19219029
捐赠科研通 7049170
什么是DOI,文献DOI怎么找? 3234644
关于科研通互助平台的介绍 2397613
邀请新用户注册赠送积分活动 2216775