计算机科学
可靠性(半导体)
计算
推论
机器学习
置信区间
人工智能
语言模型
航程(航空)
数据挖掘
区间(图论)
点估计
统计
语言理解
工作(物理)
性能预测
实证研究
统计推断
估计
度量(数据仓库)
校准
作者
Jesús M. Fraile-Hernández,Anselmo Peñas
出处
期刊:Information
[Multidisciplinary Digital Publishing Institute]
日期:2025-09-20
卷期号:16 (9): 817-817
摘要
Measuring the reliability of performance evaluations is particularly important when we evaluate non-deterministic models. This is the case of using large language models (LLMs) in classification tasks, where different runs generate different outputs. This fact raises the question about how reliable the evaluation of a solution is. Previous work relies on executing several runs and then taking some kind of average together with confidence intervals. However, confidence intervals themselves may not be reliable if the number of executions is not large enough. Therefore, more effective and robust methods are needed for their estimation. In this work, we propose a methodology that estimates model performance while capturing the intra-run variability by leveraging instance-level predictions across multiple runs, enabling the computation of more reliable confidence intervals when the gold standard is available. Our method also offers greater computational efficiency by reducing the number of full model executions required to estimate performance variability. Compared against existing state-of-the-art evaluation methods, our approach achieves full empirical coverage (100%) of plausible performance outcomes using as few as three runs, whereas traditional methods reach at most 63% coverage, even with eight runs.
科研通智能强力驱动
Strongly Powered by AbleSci AI