抽象
计算机科学
元数据
病历
编码(集合论)
医学
肌营养不良
物理医学与康复
数据收集
数据科学
数据挖掘
数据存取
数据库
梅德林
医疗急救
电子病历
情报检索
人工智能
锁(火器)
作者
Huixue Zhou,Geetanjali Rajamani,Jiatan Huang,Magali Jorand‐Fletcher,Yara Mohamed,Kody A. DeGolier,Annette Xenopoulos-Oddsson,Erjia Cui,Carla D. Zingariello,Rui Zhang,Peter B. Kang
出处
期刊:Neurology
[Lippincott Williams & Wilkins]
日期:2025-09-23
卷期号:15 (6): e200542-e200542
被引量:4
标识
DOI:10.1212/cpj.0000000000200542
摘要
Background and Objectives: Muscular dystrophies are characterized by progressive muscle weakness and degeneration. Identifying cases and abstracting data from electronic medical records (EMRs) is helpful for surveillance and research. However, manual EMR abstraction is laborious. We studied 2 approaches to accelerate EMR abstraction: large language models (LLMs) and International Classification of Diseases (ICD) code meta-analysis. Methods: criteria. Results: IAA for manual annotations varied between 80% (for annotation of symptoms) and 100% (for CK values). The highest performing LLM was Llama 3-8b, which yielded the following accuracies: 46.8% for "first symptoms," 56.9% for "ambulatory status," 69.2% for "CK values," and 68.4% for "genetic test results." Among 77 individuals with DMD, all patients with 20 or more encounters linked to relevant ICD codes had definite or probable diagnoses, whereas among 59 individuals with LGMD, all patients with 25 or more encounters linked to relevant ICD codes had definite or probable diagnoses. Discussion: LLMs promise to accelerate EMR abstraction for rare diseases such as muscular dystrophy, but F1 scores for LLMs currently lag manual abstractions for unstructured data. Llama 3-8b demonstrated superior performance to the 4 other models tested. Metadata such as ICD code counts may help prioritize high-yield cases for surveillance and research purposes.
科研通智能强力驱动
Strongly Powered by AbleSci AI