肺癌
医学
生成语法
医学物理学
放射科
计算机科学
人工智能
内科学
作者
Li‐Hua Huang,Anqi Lin,Haitao Li,Qun Wang,Junyi Shen,Aimin Jiang,Qi Chang,Wenyi Gan,Lingxuan Zhu,Weiming Mou,Dongqiang Zeng,Bufu Tang,Mingjia Xiao,Guangdi Chu,Jian Zhang,Quan Cheng,Peng Luo,Ting Wei
出处
期刊:View
[Wiley]
日期:2025-05-29
卷期号:6 (5)
被引量:2
摘要
Abstract Introduction: The emerging generative artificial intelligence (Gen‐AI) is increasingly recognized for its potential in healthcare, particularly in complex radiological interpretations. However, the clinical utility of Gen‐AI requires thorough validation with real‐world data. Method: This retrospective study analyzed chest computed tomography (CT) scans from 404 patients with lung conditions with lung neoplasms ( n = 184) and non‐malignancy ( n = 210), incorporating The Cancer Genome Atlas ( n = 106) and Medical Imaging and Data Resource Center ( n = 110) datasets as external validation. We evaluated diagnostic performance of three Gen‐AI models (GPT‐4‐turbo, Gemini‐pro‐vision, and Claude‐3‐opus) using receiver operating characteristic (ROC) analysis and chi‐square tests across various clinical scenarios. Likert scale scoring combined with response rate and variance analysis were employed to evaluate internal diagnostic tendencies, while Lasso and stepwise regression were externally introduced to optimize model performance. Results: In single‐image CT diagnostics, Gemini and Claude demonstrated superior accuracy compared to GPT. However, when additional CT slices or clinical histories were incorporated, the diagnostic accuracy of all models declined. ROC analysis indicated that Gen‐AI performance was limited but improved in simplified prompting environments or integration with machine learning methods. Feature analysis revealed that Gen‐AI primarily relied on morphology and margins for malignancy predictions, but struggled to recognize critical imaging features and occasionally fabricated data. Conclusions: Gen‐AI demonstrated variable potential for pulmonary CT imaging diagnosis across prompts and diagnostic environments of differing complexity. However, their limitations and risks in processing complex multimodal information highlight significant challenges in the integration of clinical information by existing models. Ongoing efforts to improve the robustness and reliability of these models are crucial for their successful adoption in healthcare.
科研通智能强力驱动
Strongly Powered by AbleSci AI