作者
Anirban M. Thakur,Brigid S. Mumford,P.L. Rosenblatt,Mallika Anand,K. Hanaway,William D. Winkelman
摘要
INTRODUCTION: The advent of large language models (LLMs), including Google Gemini and ChatGPT, provides patients with an accessible source of information regarding a wide variety of topics. Several studies have evaluated the accuracy and comprehensiveness of LLM responses to specialty-specific questions. To our knowledge, there are no studies that look at Google Gemini’s proficiency in responding to frequent questions patients may have in regard to urinary incontinence. OBJECTIVE: We aimed to assess the accuracy, comprehensiveness, potential for harm, and readability of LLM-generated responses to common patient questions regarding the epidemiology, definition, diagnosis, and treatment of urinary incontinence. METHODS: We queried the Google Gemini LLM with questions regarding urinary incontinence. Questions were adapted from the frequently asked questions pages of various national gynecologic and urogynecology societies and from querying Google Genesis regarding the most asked questions regarding urinary incontinence on the platform. Four board-certified urogynecology attendings analyzed each LLM-generated response on a five-point Likert Scale for 1) accuracy, 2) comprehensiveness, 3) potential for harm, and 4) readability. Questions were sorted into four categories for analysis: epidemiology, definition, diagnosis, and treatment. Five questions focused on incidence, risk factors, and natural history under epidemiology. Three questions describing urinary incontinence itself were labeled definition questions. Finally, five questions regarding workup and diagnosis were classified under diagnosis, and six questions regarding management were labeled as treatment questions (Table 1). Descriptive statistics were utilized to analyze the proportion of responses in each score, both by question category and overall. RESULTS: Among the 19 questions that were posed, composite scores were 3.9±1.0 for accuracy, 4.1±1.2 for potential for harm, 4.4±1.0 for readability. Comprehensiveness ratings were the lowest overall with an average of 3.3±1.1 (Table 2). Responses regarding epidemiology had higher ratings with accuracy noted to be 4.1±0.9 and readability 4.8±0.5. Questions pertaining to treatment had the lowest ratings, particularly in terms of accuracy (3.7±1.0) and comprehensiveness (3.0±0.09) (Table 2). Notably, 89.4% (n=17) responses recommended consultation with a healthcare provider. CONCLUSIONS: We found that Google Gemini produced responses to common patient questions regarding urinary incontinence that had high readability and low potential for harm. While Google Gemini seems to provide valuable information, reassuringly it often refers individuals to a healthcare provider. Similar to some other studies of LLM, we found that Google Gemini performed less accurately in terms of diagnosis and treatment compared to questions regarding epidemiology and those defining incontinence.