计算机科学
编码(社会科学)
自然语言处理
情报检索
人工智能
数学
统计
作者
Eyal Klang,Idit Tessler,Donald U. Apakama,Ethan Abbott,Benjamin S. Glicksberg,Monique Arnold,Akini Moses,Ankit Sakhuja,Ali Soroush,Alexander W. Charney,David L. Reich,Jolion McGreevy,Nicholas Gavin,Brendan G. Carr,Robert Freeman,Girish N. Nadkarni
出处
期刊:
日期:2025-09-25
卷期号:2 (10)
摘要
Accurate medical coding is vital for clinical and administrative purposes, but it is often complex and time-consuming. Large language models (LLMs) often struggle with medical coding, producing inaccuracies and hallucinations. We aimed to enhance LLM medical coding by integrating retrieval-augmented generation (RAG). We studied 500 randomly selected emergency department visits from the Mount Sinai Health System. Nine LLMs, both commercial and open-source, were evaluated for primary diagnosis coding according to the International Classification of Diseases, Tenth Revision, Clinical Modification. A RAG system enhanced LLM predictions using data from over 1 million emergency department visits. We compared RAG-enhanced codes with provider-assigned codes. A masked review by four physicians and two LLMs determined which codes — LLM- or provider-assigned — were more accurate and specific. RAG-enhanced LLMs demonstrated superior accuracy and specificity compared with provider-assigned codes. Human reviewers favored RAG-enhanced Generative Pretrained Transformer 4 (GPT-4) for accuracy in 447 instances, versus 277 instances for provider-assigned codes (P<0.001). For specificity, RAG-enhanced GPT-4 was preferred in 509 cases, compared with 181 for provider-assigned codes (P<0.001). Smaller open-access models also showed significant improvement with RAG, demonstrating that integrating RAG with LLM medical coding may reduce errors and improve clinical documentation. (Funded by the National Center for Advancing Translational Sciences and the Office of Research Infrastructure Programs of the National Institutes of Health.)
科研通智能强力驱动
Strongly Powered by AbleSci AI