Quality of Answers of Generative Large Language Models vs Peer Patients for Interpreting Lab Test Results for Lay Patients: Evaluation Study.

有用性 下垂 考试(生物学) 相关性(法律) 复杂度 背景(考古学) 质量(理念) 计算机科学 正确性 心理学 相似性(几何) 应用心理学 人工智能 医学教育 医学 社会心理学 政治学 古生物学 考古 社会学 程序设计语言 法学 哲学 图像(数学) 认识论 历史 生物 社会科学
作者
Zhe He,Balu Bhasuran,Qiao Jin,Shubo Tian,Karim Hanna,Cindy Shavor,Lisbeth Garcia Arguello,Patrick R. Murray,Zhiyong Lu
出处
期刊:PubMed [National Institutes of Health]
被引量:2
标识
DOI:10.2196/56655
摘要

Background: Even though patients have easy access to their electronic health records and lab test results data through patient portals, lab results are often confusing and hard to understand. Many patients turn to online forums or question and answering (Q&A) sites to seek advice from their peers. However, the quality of answers from social Q&A on health-related questions varies significantly, and not all the responses are accurate or reliable. Large language models (LLMs) such as ChatGPT have opened a promising avenue for patients to get their questions answered. Objective: We aim to assess the feasibility of using LLMs to generate relevant, accurate, helpful, and unharmful responses to lab test-related questions asked by patients and to identify potential issues that can be mitigated with augmentation approaches. Methods: We first collected lab test results related question and answer data from Yahoo! Answers and selected 53 Q&A pairs for this study. Using the LangChain framework and ChatGPT web portal, we generated responses to the 53 questions from four LLMs including GPT-4, Meta LLaMA 2, MedAlpaca, and ORCA_mini. We first assessed the similarity of their answers using standard QA similarity-based evaluation metrics including ROUGE, BLEU, METEOR, BERTScore. We also utilized an LLM-based evaluator to judge whether a target model has higher quality in terms of relevance, correctness, helpfulness, and safety than the baseline model. Finally, we performed a manual evaluation with medical experts for all the responses of seven selected questions on the same four aspects. Results: Regarding the similarity of the responses from 4 LLMs, where GPT-4 output was used as the reference answer, the responses from LLaMa 2 are the most similar ones, followed by LLaMa 2, ORCA_mini, and MedAlpaca. Human answers from Yahoo data were scored lowest and thus least similar to GPT-4-generated answers. The results of Win Rate and medical expert evaluation both showed that GPT-4's responses achieved better scores than all the other LLM responses and human responses on all the four aspects (relevance, correctness, helpfulness, and safety). However, LLM responses occasionally also suffer from lack of interpretation in one's medical context, incorrect statements, and lack of references. Conclusions: By evaluating LLMs in generating responses to patients' lab test results related questions, we find that compared to other three LLMs and human answer from the Q&A website, GPT-4's responses are more accurate, helpful, relevant, and safer. However, there are cases that GPT-4 responses are inaccurate and not individualized. We identified a number of ways to improve the quality of LLM responses including prompt engineering, prompt augmentation, retrieval augmented generation, and response evaluation.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
脑洞疼应助风中向日葵采纳,获得10
1秒前
2秒前
遂愿应助淡淡易蓉采纳,获得10
2秒前
hodi发布了新的文献求助10
2秒前
3秒前
Gilly发布了新的文献求助20
5秒前
luo发布了新的文献求助10
5秒前
呆鹅喵喵完成签到,获得积分10
5秒前
DW应助蔡从安采纳,获得10
6秒前
Sapphire完成签到,获得积分10
7秒前
JC完成签到,获得积分10
7秒前
墨锦完成签到,获得积分10
8秒前
10秒前
10秒前
夜子落完成签到 ,获得积分10
10秒前
Silence完成签到,获得积分0
13秒前
13秒前
长京完成签到 ,获得积分10
13秒前
grisco发布了新的文献求助30
14秒前
pebble完成签到,获得积分0
14秒前
16秒前
16秒前
虚心沂完成签到,获得积分10
20秒前
20秒前
20秒前
z张发布了新的文献求助10
21秒前
houruibut发布了新的文献求助10
21秒前
W43完成签到,获得积分10
22秒前
虚幻靖易完成签到,获得积分10
23秒前
在水一方应助欣宇采纳,获得10
23秒前
木南发布了新的文献求助10
25秒前
一颗苹果发布了新的文献求助10
25秒前
WWW发布了新的文献求助10
25秒前
Owen应助折耳根采纳,获得10
25秒前
luo完成签到,获得积分10
26秒前
26秒前
Gaber发布了新的文献求助10
26秒前
27秒前
28秒前
打打应助houruibut采纳,获得10
28秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Rosenblum, Global Change Biology 800
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7774085
求助须知:如何正确求助?哪些是违规求助? 9316112
关于积分的说明 20349086
捐赠科研通 7359870
什么是DOI,文献DOI怎么找? 3317352
关于科研通互助平台的介绍 2465871
邀请新用户注册赠送积分活动 2332629