Evaluating Large Language Models: A Comprehensive Survey

简编 计算机科学 风险分析(工程) 业务 地理 考古 操作系统
作者
Z. J. Guo,Renren Jin,Chuang LIU,Yufei Huang,Dongquan Shi,Supryadi,Lixin Yu,Yan Liu,Jiaxuan Li,Bin Xiong,Deyi Xiong
出处
期刊:Cornell University - arXiv [Cornell University]
标识
DOI:10.48550/arxiv.2310.19736
摘要

Large language models (LLMs) have demonstrated remarkable capabilities across a broad spectrum of tasks. They have attracted significant attention and been deployed in numerous downstream applications. Nevertheless, akin to a double-edged sword, LLMs also present potential risks. They could suffer from private data leaks or yield inappropriate, harmful, or misleading content. Additionally, the rapid progress of LLMs raises concerns about the potential emergence of superintelligent systems without adequate safeguards. To effectively capitalize on LLM capacities as well as ensure their safe and beneficial development, it is critical to conduct a rigorous and comprehensive evaluation of LLMs. This survey endeavors to offer a panoramic perspective on the evaluation of LLMs. We categorize the evaluation of LLMs into three major groups: knowledge and capability evaluation, alignment evaluation and safety evaluation. In addition to the comprehensive review on the evaluation methodologies and benchmarks on these three aspects, we collate a compendium of evaluations pertaining to LLMs' performance in specialized domains, and discuss the construction of comprehensive evaluation platforms that cover LLM evaluations on capabilities, alignment, safety, and applicability. We hope that this comprehensive overview will stimulate further research interests in the evaluation of LLMs, with the ultimate goal of making evaluation serve as a cornerstone in guiding the responsible development of LLMs. We envision that this will channel their evolution into a direction that maximizes societal benefit while minimizing potential risks. A curated list of related papers has been publicly available at https://github.com/tjunlp-lab/Awesome-LLMs-Evaluation-Papers.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
12发布了新的文献求助10
刚刚
1秒前
JamesPei应助螺蛳粉不要辣采纳,获得10
1秒前
JamesPei应助奋斗土豆采纳,获得10
2秒前
678完成签到,获得积分10
3秒前
nurbiya应助Rosaline采纳,获得10
4秒前
叶青完成签到,获得积分10
5秒前
5秒前
6秒前
7秒前
7秒前
7秒前
晓风残月发布了新的文献求助10
7秒前
8秒前
8秒前
CodeCraft应助安详胜采纳,获得10
8秒前
殷勤的小鸽子完成签到,获得积分10
8秒前
王晨光发布了新的文献求助10
9秒前
五六七发布了新的文献求助10
9秒前
9秒前
9秒前
9秒前
平淡初雪完成签到,获得积分0
10秒前
11秒前
Dr.lee完成签到,获得积分10
11秒前
11秒前
Ysk发布了新的文献求助10
11秒前
戚薇发布了新的文献求助10
12秒前
12秒前
13秒前
13秒前
CodeCraft应助dd采纳,获得10
13秒前
14秒前
14秒前
雪白丹雪完成签到 ,获得积分10
14秒前
14秒前
14秒前
只只完成签到,获得积分10
14秒前
wdf发布了新的文献求助10
15秒前
Lilian发布了新的文献求助10
15秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
日本現代怪異事典 副読本 700
Concise Introduction to Heritage Studies 650
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 650
Machine Learning for Asset Management and Pricing 600
Numerical analysis of the coupled atmosphere-ocean models (CAO II). II 600
Models for the coupled atmosphere and ocean 600
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7380625
求助须知:如何正确求助?哪些是违规求助? 8988173
关于积分的说明 19117522
捐赠科研通 7020228
什么是DOI,文献DOI怎么找? 3226829
关于科研通互助平台的介绍 2390042
邀请新用户注册赠送积分活动 2207755