Assessing the Accuracy and Readability of Generative Artificial Intelligence Responses for Esophageal and Gastric Cancer Patients

可读性 医学 利克特量表 阅读(过程) 癌症 人工智能 医学物理学 食管癌 精密医学 自然语言处理 梅德林 比例(比率) 等级 诊断准确性 机器学习 疾病 生成模型 健康素养
作者
Shanhu Ran,Wenlong Guan,Ran Wei,Yukun Chen,B L Zhang,Yating Wang,M F Zhang,Zixian Wang,Wei Liao,F Chen
出处
期刊:Journal of Clinical Medicine [Multidisciplinary Digital Publishing Institute]
卷期号:15 (8): 2958-2958
标识
DOI:10.3390/jcm15082958
摘要

Background: Generative artificial intelligence (GenAI) models are increasingly used for medical information retrieval, due to their accessibility and efficiency. However, the accuracy and readability of their responses, specifically for upper gastrointestinal cancers, remain inadequately evaluated. This gap highlights the need for rigorous assessment to ensure reliable patient education and clinical integration. Objective: This study aimed to assess the accuracy and readability of responses generated by four prominent GenAI models (Kimi, DeepSeek, ChatGPT, and Gemini) when addressing patient-focused questions related to esophageal and gastric cancers. Methods: Twenty-five standardized medical questions about esophageal and gastric cancer covering domains of disease definition, treatment and management were posed to each model. Responses were assessed by four oncologists for accuracy by a 5-point Likert scale and analyzed for readability using Flesch–Kincaid Reading Ease, Flesch–Kincaid Grade Level, and SMOG metrics. High-interest questions for patients were identified via questionnaires. Results: Comparing the accuracy of GenAI-generated responses, DeepSeek achieved the highest overall accuracy score and outperformed other models in questions about definitions and treatments, while ChatGPT excelled in management-related inquiries. In subgroup analysis, GenAI models exhibited higher accuracy in answering definition and management questions, which patients preferred to inquire, compared with questions about cancer therapies. The responses produced by all models required a reading capacity from 11th-grade to college level. Conclusions: This study revealed that in this comparative evaluation application of GenAI models, DeepSeek provides the most accurate responses for upper GI cancer inquiries about definition and treatment, while ChatGPT showed superiority in management-related questions. However, all models generate texts requiring advanced reading levels, highlighting a need for readability optimization without compromising accuracy. GenAI shows promise for patient education but requires rigorous validation for clinical integration.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
活着完成签到 ,获得积分10
刚刚
冯不疯完成签到,获得积分10
刚刚
空青完成签到,获得积分10
1秒前
1秒前
FYJ发布了新的文献求助10
1秒前
1秒前
阿高完成签到,获得积分20
1秒前
1秒前
earthclean发布了新的文献求助10
2秒前
开放的从菡完成签到 ,获得积分10
3秒前
3秒前
Shawn完成签到,获得积分10
3秒前
jiedaocheng完成签到,获得积分10
3秒前
学术文献互助应助蜀安采纳,获得200
4秒前
fairy发布了新的文献求助10
4秒前
顾矜应助小张采纳,获得10
4秒前
落幕熊猫完成签到,获得积分0
4秒前
江上浩月完成签到,获得积分10
5秒前
陶醉的大炮完成签到,获得积分10
5秒前
诚心的志泽完成签到,获得积分10
5秒前
5秒前
文艺的熠彤完成签到,获得积分10
6秒前
呱呱完成签到,获得积分10
6秒前
空青发布了新的文献求助10
6秒前
认真芷容完成签到,获得积分10
6秒前
Honey发布了新的文献求助10
6秒前
7秒前
keten完成签到,获得积分10
7秒前
7秒前
7秒前
7秒前
7秒前
7秒前
riccixuu完成签到 ,获得积分10
7秒前
8秒前
twss发布了新的文献求助10
8秒前
李皮皮完成签到 ,获得积分10
8秒前
orixero应助PositiveJugend采纳,获得10
8秒前
Weizhuo完成签到 ,获得积分10
8秒前
香蕉觅云应助落后的寻凝采纳,获得10
8秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
The anomeric effect 1000
Principles of town planning: translating concepts to applications 1000
Navigating Normative Orders: Interdisciplinary Perspectives 750
1 Peter and Christ's Descent to the Dead in Its Early Christian Reception 700
Organizational Behavior 510
Management and the Arts 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7733078
求助须知:如何正确求助?哪些是违规求助? 9283945
关于积分的说明 20161671
捐赠科研通 7310903
什么是DOI,文献DOI怎么找? 3304251
关于科研通互助平台的介绍 2457078
邀请新用户注册赠送积分活动 2313480