Evaluating Large Language Models: A Comprehensive Survey

简编 计算机科学 风险分析(工程) 业务 地理 考古 操作系统
作者
Z. J. Guo,Renren Jin,Chuang LIU,Yufei Huang,Dongquan Shi,Supryadi,Lixin Yu,Yan Liu,Jiaxuan Li,Bin Xiong,Deyi Xiong
出处
期刊:Cornell University - arXiv [Cornell University]
标识
DOI:10.48550/arxiv.2310.19736
摘要

Large language models (LLMs) have demonstrated remarkable capabilities across a broad spectrum of tasks. They have attracted significant attention and been deployed in numerous downstream applications. Nevertheless, akin to a double-edged sword, LLMs also present potential risks. They could suffer from private data leaks or yield inappropriate, harmful, or misleading content. Additionally, the rapid progress of LLMs raises concerns about the potential emergence of superintelligent systems without adequate safeguards. To effectively capitalize on LLM capacities as well as ensure their safe and beneficial development, it is critical to conduct a rigorous and comprehensive evaluation of LLMs. This survey endeavors to offer a panoramic perspective on the evaluation of LLMs. We categorize the evaluation of LLMs into three major groups: knowledge and capability evaluation, alignment evaluation and safety evaluation. In addition to the comprehensive review on the evaluation methodologies and benchmarks on these three aspects, we collate a compendium of evaluations pertaining to LLMs' performance in specialized domains, and discuss the construction of comprehensive evaluation platforms that cover LLM evaluations on capabilities, alignment, safety, and applicability. We hope that this comprehensive overview will stimulate further research interests in the evaluation of LLMs, with the ultimate goal of making evaluation serve as a cornerstone in guiding the responsible development of LLMs. We envision that this will channel their evolution into a direction that maximizes societal benefit while minimizing potential risks. A curated list of related papers has been publicly available at https://github.com/tjunlp-lab/Awesome-LLMs-Evaluation-Papers.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
三三发布了新的文献求助30
1秒前
cao完成签到,获得积分10
1秒前
2秒前
充电宝应助张静采纳,获得10
2秒前
inm323完成签到,获得积分10
2秒前
搜集达人应助Hear采纳,获得10
3秒前
3秒前
3秒前
菲露詹发布了新的文献求助10
4秒前
田乐天完成签到 ,获得积分10
6秒前
小马甲应助zspu163采纳,获得10
6秒前
6秒前
6秒前
7秒前
Circle发布了新的文献求助10
9秒前
9秒前
megac完成签到,获得积分10
9秒前
9秒前
Nini1203发布了新的文献求助10
10秒前
痴情的凌兰完成签到,获得积分10
10秒前
时尚蓝发布了新的文献求助10
11秒前
11秒前
杜明智发布了新的文献求助10
12秒前
12秒前
小李完成签到,获得积分10
16秒前
16秒前
16秒前
17秒前
17秒前
17秒前
17秒前
孔祥柏完成签到,获得积分10
17秒前
科研通AI6.2应助shane采纳,获得10
18秒前
冷如松发布了新的文献求助10
18秒前
沉静秋尽完成签到,获得积分10
18秒前
华仔应助小李采纳,获得10
19秒前
19秒前
小蒋完成签到,获得积分10
20秒前
顾矜应助yiming采纳,获得10
20秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Geist der Kunst und Kultur 1000
Resistance Spot Welding Dataset for Automobile Body-in-White Quality Analysis 748
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 700
Machine Learning for Asset Management and Pricing 600
Numerical analysis of the coupled atmosphere-ocean models (CAO II). II 600
Models for the coupled atmosphere and ocean 600
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7406474
求助须知:如何正确求助?哪些是违规求助? 9010873
关于积分的说明 19190566
捐赠科研通 7039828
什么是DOI,文献DOI怎么找? 3232337
关于科研通互助平台的介绍 2394360
邀请新用户注册赠送积分活动 2214477