已入深夜,您辛苦了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!祝你早点完成任务,早点休息,好梦!

Evaluating large language models for software testing

计算机科学 软件测试 软件工程 程序设计语言 软件
作者
Yihao Li,Pan Liu,Haiyang Wang,Jie Chu,W. Eric Wong
出处
期刊:Computer Standards & Interfaces [Elsevier BV]
卷期号:93: 103942-103942 被引量:29
标识
DOI:10.1016/j.csi.2024.103942
摘要

• We present a new evaluation framework for testing the capabilities of LLMs , using manual testing as a benchmark to assess different LLMs. This approach avoids the potential reliance on existing test sets that may be part of LLM training data , ensuring a more reliable evaluation of their generalization capabilities. • We utilize third-party open-source software for LLM testing evaluation, ensuring bug reproducibility and simulating the real-world application of LLMs in software testing, thus enabling more objective evaluations. • We evaluate different LLMs from multiple perspectives and discuss practical strategies for implementing LLM-driven testing effectively. • We propose the follow-up question method for LLM-driven testing, which shows promise in enhancing the abilities of LLMs to detect bugs and defects in program code. Large language models (LLMs) have demonstrated significant prowess in code analysis and natural language processing, making them highly valuable for software testing. This paper conducts a comprehensive evaluation of LLMs applied to software testing, with a particular emphasis on test case generation, error tracing, and bug localization across twelve open-source projects. The advantages and limitations, as well as recommendations associated with utilizing LLMs for these tasks, are delineated. Furthermore, we delve into the phenomenon of hallucination in LLMs, examining its impact on software testing processes and presenting solutions to mitigate its effects. The findings of this work contribute to a deeper understanding of integrating LLMs into software testing, providing insights that pave the way for enhanced effectiveness in the field.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
tcxty完成签到,获得积分10
刚刚
hhhjy发布了新的文献求助10
1秒前
Nole应助科研通管家采纳,获得10
1秒前
1秒前
小Q发布了新的文献求助10
1秒前
1秒前
Nole应助科研通管家采纳,获得10
1秒前
脑洞疼应助科研通管家采纳,获得10
2秒前
2秒前
Nole应助科研通管家采纳,获得10
2秒前
2秒前
王博士完成签到,获得积分10
4秒前
甜美的秋尽完成签到,获得积分10
5秒前
英俊的铭应助weilan采纳,获得10
5秒前
5秒前
NexusExplorer应助Ming采纳,获得50
5秒前
6秒前
祝福发布了新的文献求助10
6秒前
北风语完成签到,获得积分10
7秒前
8秒前
zzz完成签到 ,获得积分10
8秒前
33499083发布了新的文献求助10
9秒前
可爱的函函应助正直未来采纳,获得10
9秒前
10秒前
烟花应助小Q采纳,获得10
11秒前
KID发布了新的文献求助10
11秒前
小何完成签到 ,获得积分10
11秒前
12秒前
11发布了新的文献求助10
12秒前
久热发布了新的文献求助10
12秒前
科研通AI6.2应助壮观若南采纳,获得10
14秒前
15秒前
FashionBoy应助33499083采纳,获得10
15秒前
晚意意意意意完成签到 ,获得积分10
16秒前
ZMTW发布了新的文献求助10
16秒前
顾矜应助雪山冰川采纳,获得10
16秒前
Yebb完成签到,获得积分10
16秒前
17秒前
忧郁发卡发布了新的文献求助10
20秒前
万能图书馆应助yu采纳,获得30
21秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Rosenblum, Global Change Biology 800
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Physiologic specialization in Peronospora manshurica 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7777672
求助须知:如何正确求助?哪些是违规求助? 9318513
关于积分的说明 20364691
捐赠科研通 7364587
什么是DOI,文献DOI怎么找? 3318990
关于科研通互助平台的介绍 2466628
邀请新用户注册赠送积分活动 2334211