作者
Chiaki Yuasa,Ryunosuke Noda,Fumiya Kitano,Daisuke Ichikawa,Yugo Shibagaki
摘要
Abstract Background and Aims Large language models (LLMs) such as GPT-4 have demonstrated strong performance on medical exams, but their capacities to handle specialized domains like nephrology remain under investigation. Recently, a next-generation LLM called o1 pro was released, claiming superior reasoning abilities and fewer hallucinations. This study evaluated whether o1 pro could outperform GPT-4 on a rigorous set of Japanese nephrology board renewal questions. By comparing their accuracy and consistency across various subdomains and question types, we aimed to clarify the potential of these models as clinical and educational tools in nephrology. Method We used the Self-Assessment Questions for Nephrology Board Renewal, a set of annually administered, Japanese-language multiple-choice questions. These 209 questions (covering 2014–2023) encompass fundamental concepts, clinical management, and image interpretation in nephrology. Each question was categorized by cognitive level (recall, interpretation, and problem-solving), question type (general vs. clinical), the presence or absence of images, and nephrology subspecialty. Both o1 pro and GPT-4 were tested through separate chat sessions for each item, with accuracy determined by official answers from the Japanese Society of Nephrology. Statistical comparisons of correct-answer proportions were made using chi-square or Fisher's exact tests. A p-value of less than 0.05 was considered significant. Results Overall, o1 pro achieved a proportion of correct answers of 81.3% (170/209), significantly exceeding GPT-4’s 51.2% (107/209; P < 0.001). When analyzed by examination year, o1 pro consistently surpassed the 60% pass threshold and exhibited strong annual performance (70%–95% accuracy). In contrast, GPT-4 displayed notable variability and passed the 60% mark in only two of the ten examination years. By cognitive level, o1 pro showed superior accuracy for recall (83.3% vs. 49.1%), interpretation (75.0% vs. 50.0%), and problem-solving (84.4% vs. 57.8%) questions (P < 0.001, 0.011, 0.011, respectively). Similar trends were seen for general versus clinical question types: o1 pro achieved 83.8% and 78.6%, whereas GPT-4 attained 49.5% and 53.1% (P < 0.001, < 0.001, respectively). Additionally, o1 pro displayed marked advantages on image-based questions (77.5% vs. 42.5%; P = 0.003), suggesting enhanced capacities in multimodal information processing. Subspecialty analyses confirmed that o1 pro outperformed GPT-4 in critical areas like chronic kidney disease/end-stage kidney disease. Conclusion This study demonstrated that o1 pro significantly outperformed GPT-4 in multiple facets of nephrology board renewal questions, spanning basic recall, complex clinical reasoning, and image-based interpretation. These findings underscore the rapidly evolving capabilities of LLMs and their potential as educational and clinical support tools in nephrology. As further refinements in LLM architecture, training data, and multimodal integration continue, o1 pro and similarly advanced models may play an increasingly vital role in enhancing nephrology education and decision-making. However, real-world validation and assessments across diverse medical settings are necessary before fully integrating next-generation LLMs into routine clinical practice.