医学
组内相关
射线照相术
分级(工程)
骨关节炎
骨科手术
等级间信度
口腔正畸科
物理疗法
放射科
外科
评定量表
统计
病理
心理测量学
数学
土木工程
替代医学
工程类
临床心理学
作者
Jiesheng Zhu,Yinuo Jiang,Daosen Chen,Yi Lu,Yijiang Huang,Yimu Lin,Pei Fan
摘要
PURPOSE: To explore the potential of ChatGPT-4o in analysing radiographic images of knee osteoarthritis (OA) and to assess its grading accuracy, feature identification and reliability, thereby helping surgeons to improve diagnostic accuracy and efficiency. METHODS: A total of 117 anterior‒posterior knee radiographs from patients (23.1% men, 76.9% women, mean age 69.7 ± 7.99 years) were analysed. Two senior orthopaedic surgeons and ChatGPT-4o independently graded images with the Kellgren-Lawrence (K-L), Ahlbäck and International Knee Documentation Committee (IKDC) systems. A consensus reference standard was established by a third radiologist. ChatGPT-4o's performance metrics (accuracy, precision, recall and F1 score) were calculated, and its reliability was assessed via two evaluations separated by a 2-week interval, with intraclass correlation coefficients (ICCs) determined. RESULTS: ChatGPT-4o achieved a 100% identification rate for knee radiographs and demonstrated strong binary classification performance (precision: 0.95, recall: 0.83, F score: 0.88). However, its detailed grading accuracy (35%) was substantially lower than that of surgeons (89.6%). Severe underestimation of OA severity occurred in 49.3% of the cases. Interrater reliability for surgeons was excellent (ICC: 0.78-0.91), whereas ChatGPT-4o showed poor initial consistency (ICC: 0.16-0.28), improving marginally in the second evaluation (ICC: 0.22-0.39). CONCLUSION: ChatGPT-4o has the potential to rapidly identify and binary classify knee OA on radiographs. However, its detailed grading accuracy remains suboptimal, with a notable tendency to underestimate severe cases. This limits its current clinical utility for precise staging. Future research should focus on optimising its grading performance and improving accuracy to enhance diagnostic reliability. LEVEL OF EVIDENCE: Level III, retrospective comparative study.
科研通智能强力驱动
Strongly Powered by AbleSci AI