On Measuring Large Language Models Performance with Inferential Statistics

计算机科学 可靠性(半导体) 计算 推论 机器学习 置信区间 人工智能 语言模型 航程(航空) 数据挖掘 区间(图论) 点估计 统计 语言理解 工作(物理) 性能预测 实证研究 统计推断 估计 度量(数据仓库) 校准
作者
Jesús M. Fraile-Hernández,Anselmo Peñas
出处
期刊:Information [Multidisciplinary Digital Publishing Institute]
卷期号:16 (9): 817-817
标识
DOI:10.3390/info16090817
摘要

Measuring the reliability of performance evaluations is particularly important when we evaluate non-deterministic models. This is the case of using large language models (LLMs) in classification tasks, where different runs generate different outputs. This fact raises the question about how reliable the evaluation of a solution is. Previous work relies on executing several runs and then taking some kind of average together with confidence intervals. However, confidence intervals themselves may not be reliable if the number of executions is not large enough. Therefore, more effective and robust methods are needed for their estimation. In this work, we propose a methodology that estimates model performance while capturing the intra-run variability by leveraging instance-level predictions across multiple runs, enabling the computation of more reliable confidence intervals when the gold standard is available. Our method also offers greater computational efficiency by reducing the number of full model executions required to estimate performance variability. Compared against existing state-of-the-art evaluation methods, our approach achieves full empirical coverage (100%) of plausible performance outcomes using as few as three runs, whereas traditional methods reach at most 63% coverage, even with eight runs.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
xc关闭了xc文献求助
5秒前
5秒前
充电宝应助酷炫的雪枫采纳,获得10
5秒前
yeweijia完成签到,获得积分10
6秒前
6秒前
隐形曼青应助luxinyue采纳,获得10
6秒前
林阙应助炽热的光采纳,获得30
6秒前
fxy发布了新的文献求助10
7秒前
7秒前
张思成发布了新的文献求助10
7秒前
黄诗婷完成签到,获得积分10
8秒前
8秒前
8秒前
烟鹤发布了新的文献求助10
8秒前
8秒前
东方不败完成签到 ,获得积分10
9秒前
9秒前
9秒前
MMY发布了新的文献求助30
10秒前
乐乐应助Benthesikyme采纳,获得10
10秒前
11秒前
谢雷XIELei应助西法豆腐采纳,获得10
11秒前
11秒前
结实的小馒头完成签到,获得积分10
12秒前
1113发布了新的文献求助10
12秒前
12秒前
13秒前
科研通AI6.4应助张思成采纳,获得10
13秒前
13秒前
爱笑麦丽素完成签到 ,获得积分10
13秒前
Owen应助成就的冬卉采纳,获得10
13秒前
14秒前
14秒前
Jared发布了新的文献求助10
15秒前
深情安青应助liushu采纳,获得10
15秒前
15秒前
16秒前
hrs发布了新的文献求助10
16秒前
汉堡包应助kkkk采纳,获得10
16秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
The Multiple Self-States Drawing Technique 600
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Rosenblum, Global Change Biology 500
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7770586
求助须知:如何正确求助?哪些是违规求助? 9313487
关于积分的说明 20334126
捐赠科研通 7356011
什么是DOI,文献DOI怎么找? 3316465
关于科研通互助平台的介绍 2465126
邀请新用户注册赠送积分活动 2331293