作者
Lu Wang,Yuqiang Mao,Lin Zhang,Jiangdian Song
摘要
Background: Large language models are proficient in text- and image-based medical tasks; however, their performance in video-based tasks remains unclear. We evaluated the performance of the state-of-the-art GPT-4o in lung nodule feature characterization and lung cancer probability estimation for computed tomography (CT) images of lungs using both China and the open-access LIDC-IDRI datasets. Methods: Study design: Participants with CT scans performed between March 31, 2023, and March 31, 2024, were retrieved from two centers (represented by D1 and D2) in China. Patients in D1 and D2 were required with visible lung lesions and pathological result. For the LIDC-IDRI dataset, this study included participants in which four radiologists from the LIDC-IDRI program annotated the location, size, and likelihood of malignancy of the lung nodules. To further evaluate GPT-4o’s performance for lung cancer screening, we retrieved data from the first in-house center to enroll participants with low-dose CT scans to form the D3 dataset. Pathological results were not required for D3. This retrospective study was approved by our institutional review board.Participants: A total of 795 participants were enrolled, including 240 in D1 (with 120 malignancy and 120 benign cases), 205 in D2 (with 104 malignancy and 101 benign cases), 150 in D3, and 200 in LIDC-IDRI dataset.Analysis: After anonymization, CT images of each participant were converted into a video using 12 images per second (arranged sequentially from the neck to the abdomen). This study first used the data of 20 participants to create textual prompts for GPT-4o. For lung nodule recognition, the prompt was provided with nodule coordinate and then GPT-4o automatically recognizing the nodule within the CT scan and generating an image with a bounding box around its identified nodule. Two radiologists independently assessed whether GPT-4o successfully identified the lung nodule based on the provided bounding box image. For the lung cancer estimation task on the D1 and D2 datasets, the quantitative probability of malignancy (ranging from 1% to 100%) was required in the GPT-4o’s output. GPT-4o was asked to provide features of morphologies, margins, internal structures, enhancements, and surrounding structural characteristics of each nodule. Six radiologists characterized the lung nodule features and then graded the agreement with GPT-4o’s reports on a 5-point Likert scale (1=completely incorrect and 5=completely correct). GPT-4o’s malignancy probability estimations for lung cancer screening were compared to two radiologists’ Lung-RADS scores on D3. Findings: GPT-4o achieved accuracies of 95.2%, 93.4%, and 95.5% for lung nodule region recognition for the D1, D2, and LIDC-IDRI datasets, respectively. Compared to pathology results, the AUC of GPT-4o’s probability was 0.70 (95% CI: 0.64–0.78) and 0.73 (95% CI: 0.66–0.80) on D1 and D2. The radiologists’ evaluations confirmed GPT-4o’s performance with median agreements of 4.1 (IQR, 3.2–4.5) and 4.2 (IQR, 3.5–4.7) on D1 and D2, and an intraclass correlation coefficient (ICC) of 0.79 on the LIDC-IDRI dataset, and an ICC of 0.60 on D3 for lung cancer screening. Interpretation: The results showed that GPT-4o is promising for assisting lung nodule feature characterization and lung cancer probability estimation in radiology.