An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation

计算机科学 单元测试 正确性 文档 JavaScript 考试(生物学) 自动化 人工智能 软件 软件工程 机器学习 程序设计语言 机械工程 生物 工程类 古生物学
作者
Max Schäfer,Sarah Nadi,Aryaz Eghbali,Frank Tip
出处
期刊:IEEE Transactions on Software Engineering [IEEE Computer Society]
卷期号:50 (1): 85-105 被引量:264
标识
DOI:10.1109/tse.2023.3334955
摘要

Unit tests play a key role in ensuring the correctness of software. However, manually creating unit tests is a laborious task, motivating the need for automation. Large Language Models (LLMs) have recently been applied to various aspects of software development, including their suggested use for automated generation of unit tests, but while requiring additional training or few-shot learning on examples of existing tests. This paper presents a large-scale empirical evaluation on the effectiveness of LLMs for automated unit test generation without requiring additional training or manual effort. Concretely, we consider an approach where the LLM is provided with prompts that include the signature and implementation of a function under test, along with usage examples extracted from documentation. Furthermore, if a generated test fails, our approach attempts to generate a new test that fixes the problem by re-prompting the model with the failing test and error message. We implement our approach in TestPilot , an adaptive LLM-based test generation tool for JavaScript that automatically generates unit tests for the methods in a given project's API. We evaluate TestPilot using OpenAI's gpt3.5-turbo LLM on 25 npm packages with a total of 1,684 API functions. The generated tests achieve a median statement coverage of 70.2% and branch coverage of 52.8%. In contrast, the state-of-the feedback-directed JavaScript test generation technique, Nessie, achieves only 51.3% statement coverage and 25.6% branch coverage. Furthermore, experiments with excluding parts of the information included in the prompts show that all components contribute towards the generation of effective test suites. We also find that 92.8% of TestPilot 's generated tests have $\leq$ 50% similarity with existing tests (as measured by normalized edit distance), with none of them being exact copies. Finally, we run TestPilot with two additional LLMs, OpenAI's older code-cushman-002 LLM and StarCoder , an LLM for which the training process is publicly documented. Overall, we observed similar results with the former (68.2% median statement coverage), and somewhat worse results with the latter (54.0% median statement coverage), suggesting that the effectiveness of the approach is influenced by the size and training set of the LLM, but does not fundamentally depend on the specific model.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
克拉发布了新的文献求助10
1秒前
1秒前
谢言一完成签到,获得积分10
1秒前
体贴的醉山完成签到,获得积分10
1秒前
li发布了新的文献求助10
1秒前
彭于晏应助DuanJN采纳,获得10
1秒前
1秒前
1秒前
rhea完成签到 ,获得积分10
2秒前
lcyxdsl完成签到,获得积分10
2秒前
周小鱼完成签到,获得积分10
2秒前
葛二蛋发布了新的文献求助20
2秒前
张张发布了新的文献求助10
2秒前
PLUS完成签到,获得积分20
2秒前
李晶晶完成签到 ,获得积分20
2秒前
3秒前
3秒前
百川完成签到,获得积分10
3秒前
大个应助菠菠柑采纳,获得10
4秒前
幼兰呆鹅发布了新的文献求助10
4秒前
沐晴完成签到,获得积分10
4秒前
柳絮发布了新的文献求助10
4秒前
4秒前
5秒前
研友_VZG7GZ应助yu采纳,获得10
5秒前
6秒前
王博林发布了新的文献求助10
6秒前
6秒前
6秒前
碧蓝亦玉完成签到,获得积分10
6秒前
renxiaoting发布了新的文献求助10
6秒前
6秒前
7秒前
7秒前
奈何发布了新的文献求助10
7秒前
123应助储存求助采纳,获得10
7秒前
光_电发布了新的文献求助10
7秒前
Tori_Q完成签到,获得积分10
7秒前
7秒前
Achilles完成签到,获得积分10
8秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Les chinois de jakarta: temples et vie collective 1000
Autoparametric Resonance in Mechanical Systems 1000
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 800
Social Psychology 600
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 600
Management and the Arts 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7646758
求助须知:如何正确求助?哪些是违规求助? 9219078
关于积分的说明 19784347
捐赠科研通 7211712
什么是DOI,文献DOI怎么找? 3277190
关于科研通互助平台的介绍 2438693
邀请新用户注册赠送积分活动 2275419