可比性
软件部署
计算机科学
医疗保健
钥匙(锁)
知识管理
数据科学
管理科学
自然语言
大数据
风险分析(工程)
医疗保健系统
专家系统
过程管理
人工智能
作者
Qing Chang,Fei Chen,Yaolong Chen,Longlong Cheng,Di Dong,Jiahong Dong,Xiaobin Feng,Junbo Ge,Jingjing He,Yihua He,Zhiyang He,Hong Ji,Xue Jiang,Zehua Jiang,Nan Li,Peng Li,Yazi Li,Bing Liu,Junwei Liu,Han Lyu
标识
DOI:10.1016/j.imed.2025.09.001
摘要
Large Language Models (LLMs), trained on vast amounts of textual data, have demonstrated strong capabilities in natural language understanding and generation. In the medical field, LLMs are increasingly applied across various domains such as disease screening, diagnostic assistance, and health management, playing a key role in advancing intelligent healthcare. In recent years, China has actively promoted the integration of artificial intelligence with healthcare through a series of policies that support enterprises in making breakthroughs in key technologies such as medical large language models and multi-modal data integration. Concurrently, efforts have accelerated the deployment of AI in applications such as health management and precision medicine, to gradually establish a full-cycle intelligent healthcare system encompassing prevention, diagnosis, treatment, and rehabilitation. However, the rapid deployment of LLMs in healthcare has highlighted the lack of standardized evaluation criteria and consistent methodologies. To address this, this expert consensus focuses on establishing a retrospective evaluation framework tailored to medical applications. By integrating scientific evaluation metrics, standards, and procedures, the framework provides clear and actionable guidance for model evaluators, developers, and end-users. It aims to unify assessment practices, enhance the scientific rigor and comparability of evaluations, and ensure the safe and effective use of LLMs in healthcare, ultimately supporting the high-quality development of AI-powered medical services.
科研通智能强力驱动
Strongly Powered by AbleSci AI