作者
Yosef Adiniaev,Mahmud Omar,Tohar M. Timor,Yiftach Barash,Olga R. Brook,Mohammad E. Naffaa,Alon Gorenshtein,Eyal Klang
摘要
OBJECTIVE: Large language models (LLMs) are increasingly evaluated for rheumatology tasks, but their performance in inflammatory arthritis remains unclear. We systematically reviewed LLM performance across clinical tasks in inflammatory arthritis. METHODS: We conducted a systematic review (PROSPERO: CRD420261359100), searching PubMed, Scopus, and PubMed Central (January 2022 to April 2026) for studies evaluating LLM performance on clinical tasks in inflammatory arthritis. Two reviewers (Y.A., A.G.) screened 113 records. RESULTS: Eighteen studies covered rheumatoid arthritis (n=3), ankylosing spondylitis/axial spondyloarthritis (n=7), psoriatic arthritis (n=2), gout (n=1), juvenile idiopathic arthritis (n=1), and multiple diseases (n=4). Most diseases and tasks were represented by only one to a few studies, and the evidence base remains earlystage and uneven across conditions. Over 20 distinct LLMs were evaluated, including ChatGPT-3.5 to ChatGPT-4o, Gemini 2.0, DeepSeek-R1/V3, Claude, and Perplexity; ChatGPT/GPT variants were the most frequently tested models (16 of 18 studies), so the current evidence base is predominantly GPT/ChatGPT-based. Findings spanned patient education (n=11), guideline adherence (n=6), clinical reasoning (n=3), and other applications (n=1). All readability assessments exceeded recommended thresholds. Guideline concordance ranged from 48% to 96%. Accuracy was lower for case-based clinical scenarios (4.24/6) than FAQ and guideline-based questions (5.32-5.36/6; p=0.044). When compared with real clinical data, agreement was poor (Cohen and Fleiss κ ≈ 0). CONCLUSION: LLMs may support patient education, factual medication queries, and structured guideline questions when used under clinician review, but should not be used for case-based reasoning, treatment selection, or autonomous clinical decisions. None of the 18 included studies evaluated retrieval-augmented or agent-based systems, and none prospectively validated LLMs in clinical workflows. Safe integration in rheumatology will require purpose-built, knowledge-grounded systems and prospective evaluation before routine clinical use.