亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

Evaluating the effectiveness of biomedical fine-tuning for large language models on clinical tasks

计算机科学 自动汇总 语言模型 编码(社会科学) 人工智能 数学 统计
作者
Felix J. Dorfner,Amin Dada,Felix Busch,Marcus R. Makowski,Tianyu Han,Daniel Truhn,Jens Kleesiek,Madhumita Sushil,Lisa C. Adams,Keno K. Bressem
出处
期刊:Journal of the American Medical Informatics Association [Oxford University Press]
卷期号:32 (6): 1015-1024 被引量:27
标识
DOI:10.1093/jamia/ocaf045
摘要

OBJECTIVES: Large language models (LLMs) have shown potential in biomedical applications, leading to efforts to fine-tune them on domain-specific data. However, the effectiveness of this approach remains unclear. This study aims to critically evaluate the performance of biomedically fine-tuned LLMs against their general-purpose counterparts across a range of clinical tasks. MATERIALS AND METHODS: We evaluated the performance of biomedically fine-tuned LLMs against their general-purpose counterparts on clinical case challenges from NEJM and JAMA, and on multiple clinical tasks, such as information extraction, document summarization and clinical coding. We used a diverse set of benchmarks specifically chosen to be outside the likely fine-tuning datasets of biomedical models, ensuring a fair assessment of generalization capabilities. RESULTS: Biomedical LLMs generally underperformed compared to general-purpose models, especially on tasks not focused on probing medical knowledge. While on the case challenges, larger biomedical and general-purpose models showed similar performance (eg, OpenBioLLM-70B: 66.4% vs Llama-3-70B-Instruct: 65% on JAMA), smaller biomedical models showed more pronounced underperformance (OpenBioLLM-8B: 30% vs Llama-3-8B-Instruct: 64.3% on NEJM). Similar trends appeared across CLUE benchmarks, with general-purpose models often achieving higher scores in text generation, question answering, and coding. Notably, biomedical LLMs also showed a higher tendency to hallucinate. DISCUSSION: Our findings challenge the assumption that biomedical fine-tuning inherently improves LLM performance, as general-purpose models consistently performed better on unseen medical tasks. Retrieval-augmented generation may offer a more effective strategy for clinical adaptation. CONCLUSION: Fine-tuning LLMs on biomedical data may not yield the anticipated benefits. Alternative approaches, such as retrieval augmentation, should be further explored for effective and reliable clinical integration of LLMs.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
Ususl发布了新的文献求助10
5秒前
仁爱的鞋子完成签到,获得积分10
9秒前
深情安青应助Ususl采纳,获得10
16秒前
Seasun完成签到,获得积分10
30秒前
Sam1357发布了新的文献求助50
33秒前
CodeCraft应助结实的抽屉采纳,获得10
35秒前
清脆雅柔完成签到,获得积分10
36秒前
GWX完成签到,获得积分10
49秒前
顺利盼曼完成签到 ,获得积分10
55秒前
55秒前
58秒前
烟花应助科研通管家采纳,获得10
59秒前
旋光完成签到 ,获得积分10
1分钟前
hu发布了新的文献求助10
1分钟前
slouchy完成签到 ,获得积分10
1分钟前
旋光关注了科研通微信公众号
1分钟前
1分钟前
鸿影发布了新的文献求助10
1分钟前
典雅依玉完成签到,获得积分10
1分钟前
Joy完成签到 ,获得积分10
1分钟前
1分钟前
1分钟前
英勇的不斜完成签到,获得积分10
1分钟前
1分钟前
纯真的雁凡完成签到,获得积分10
1分钟前
鸿影完成签到,获得积分10
1分钟前
哩哩完成签到,获得积分20
1分钟前
谦让的鹤轩完成签到,获得积分10
1分钟前
1分钟前
英勇问晴完成签到,获得积分10
2分钟前
2分钟前
深情雪珊完成签到,获得积分10
2分钟前
纯情的阁发布了新的文献求助50
2分钟前
情怀应助读书的时候采纳,获得10
2分钟前
柔弱的铅笔完成签到,获得积分10
2分钟前
田様应助科研通管家采纳,获得30
2分钟前
Owen应助科研通管家采纳,获得10
2分钟前
不安的晓露完成签到,获得积分10
3分钟前
3分钟前
liuye0202完成签到,获得积分10
3分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
The anomeric effect 1000
Principles of town planning: translating concepts to applications 1000
1 Peter and Christ's Descent to the Dead in Its Early Christian Reception 700
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7732444
求助须知:如何正确求助?哪些是违规求助? 9283150
关于积分的说明 20156337
捐赠科研通 7309794
什么是DOI,文献DOI怎么找? 3304079
关于科研通互助平台的介绍 2456847
邀请新用户注册赠送积分活动 2313162