Large Language Model Evaluation in Traditional Chinese Medicine for Stroke: Quantitative Benchmarking Study

标杆管理 计算机科学 水准点(测量) 人工智能 数据科学 中医药 实证研究 自然语言处理 机器学习 管理科学 基础(证据) 语言模型 定量分析(化学) 知识管理 中文 深度学习 中国 运筹学 数据挖掘
作者
Hulin Long,Yang Deng,Yaoguang Guo,Zifan Shen,Yuzhu Zhang,Ji Bao,Yang He
出处
期刊:JMIR formative research [JMIR Publications Inc.]
卷期号:9: e81545-e81545
标识
DOI:10.2196/81545
摘要

BACKGROUND: The application of large language models (LLMs) in medicine is rapidly advancing. However, evaluating LLM capabilities in specialized domains such as traditional Chinese medicine (TCM), which possesses a unique theoretical system and cognitive framework, remains a sizable challenge. OBJECTIVE: This study aimed to provide an empirical evaluation of different LLM types in the specialized domain of TCM stroke. METHODS: The Traditional Chinese Medicine-Stroke Evaluation Dataset (TCM-SED), a 203-question benchmark, was systematically constructed. The dataset includes 3 paradigms (short-answer questions, multiple-choice questions, and essay questions) and covers multiple knowledge dimensions, including diagnosis, pattern differentiation and treatment, herbal formulas, acupuncture, interpretation of classical texts, and patient communication. Gold standard answers were established through a multiexpert cross-validation and consensus process. The TCM-SED was subsequently used to comprehensively test 2 representative LLM models: GPT-4o (a leading international general-purpose model) and DeepSeek-R1 (a large model primarily trained on Chinese corpora). RESULTS: The test results revealed a differentiation in model capabilities across cognitive levels. In objective sections emphasizing precise knowledge recall, DeepSeek-R1 comprehensively outperformed GPT-4o, achieving an accuracy lead of more than 17% in the multiple-choice section (96/137, 70.1% vs 72/137, 52.6%, respectively). Conversely, in the essay section, which tested knowledge integration and complex reasoning, GPT-4o's performance notably surpassed that of DeepSeek-R1. For instance, in the interpretation of classical texts category, GPT-4o achieved a scoring rate of 90.5% (181/200), far exceeding DeepSeek-R1 (147/200, 73.5%). CONCLUSIONS: This empirical study demonstrates that Chinese-centric models have a substantial advantage in static knowledge tasks within the TCM domain, whereas leading general-purpose models exhibit stronger dynamic reasoning and content generation capabilities. The TCM-SED, developed as the benchmark for this study, serves as an effective quantitative tool for evaluating and selecting appropriate LLMs for TCM scenarios. It also offers a valuable data foundation and a new research direction for future model optimization and alignment.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
yvetta发布了新的文献求助10
刚刚
1秒前
1秒前
www发布了新的文献求助10
1秒前
hachii发布了新的文献求助20
2秒前
远游客发布了新的文献求助10
2秒前
cijing发布了新的文献求助10
3秒前
竹墨发布了新的文献求助20
4秒前
SKF发布了新的文献求助20
5秒前
SKF发布了新的文献求助20
5秒前
6秒前
6秒前
7秒前
aaaa应助hurricane188采纳,获得30
7秒前
春山完成签到,获得积分10
8秒前
动听千山发布了新的文献求助30
8秒前
肥肥发布了新的文献求助10
8秒前
9秒前
9秒前
10秒前
3sigma完成签到,获得积分10
10秒前
10秒前
916发布了新的文献求助10
10秒前
11秒前
916发布了新的文献求助10
11秒前
搜集达人应助XR采纳,获得10
12秒前
隐形曼青应助Dr.c采纳,获得10
12秒前
13秒前
子安完成签到 ,获得积分10
13秒前
916发布了新的文献求助10
13秒前
春山发布了新的文献求助10
14秒前
916发布了新的文献求助10
14秒前
916发布了新的文献求助10
14秒前
15秒前
15秒前
15秒前
15秒前
小川发布了新的文献求助10
16秒前
搜集达人应助hachii采纳,获得10
17秒前
916发布了新的文献求助30
17秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
A Case Study on Hotels as Noncongregate Emergency Living Accommodations for Returning Citizens 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7765000
求助须知:如何正确求助?哪些是违规求助? 9309358
关于积分的说明 20310654
捐赠科研通 7349841
什么是DOI,文献DOI怎么找? 3314708
关于科研通互助平台的介绍 2464103
邀请新用户注册赠送积分活动 2329140