指南
更安全的
心力衰竭
射血分数
人工智能
医学诊断
医学
计算机科学
临床实习
决策支持系统
临床决策支持系统
考试(生物学)
机器学习
语言模型
领域(数学分析)
基线(sea)
强化学习
风险管理
分数(化学)
梅德林
风险评估
医学物理学
F1得分
自然语言处理
医疗保健
患者安全
多序列比对
医疗急救
临床决策
作者
Lu Liu,Chenchen Dong,Yunbo Ba,Haihong Yan,Xiaoxiao Tang,Yu Sun,Huilin Chen,Boyuan Shi,Qin Yu,Shulong Zhang
出处
期刊:Digital health
[SAGE Publishing]
日期:2026-02-01
卷期号:12: 20552076261487568-20552076261487568
标识
DOI:10.1177/20552076261487568
摘要
Objectives: Large language models (LLMs) are increasingly studied for clinical decision support, but high-risk cardiology exposes persistent weaknesses in hallucination control, guideline adherence, and medication-safety reasoning. Heart failure with reduced ejection fraction (HFrEF) is a demanding test case because safe care requires structured guideline-directed therapy, comorbidity-aware monitoring, and reliable risk warnings. Methods: We developed a dynamic alignment framework using 1087 retrospective HFrEF cases from Affiliated Zhongshan Hospital of Dalian University. An open-source LLaMA-3.1 backbone was optimized through four sequential stages: continual pre-training for heart-failure domain adaptation, supervised fine-tuning for structured clinical responses, reinforcement policy optimization for safety-oriented alignment, and retrieval-augmented generation for guideline grounding. Models were assessed with dual-track clinical and linguistic metrics. Results: LLaMA-3.1 was the strongest supervised baseline, but supervised fine-tuning alone did not fully resolve guideline-adherence limitations. Staged alignment produced a measurable Alignment Tax: the final retrieval-grounded variant improved the Clinical Score from 0.716 to 0.864 and reached a Guideline Score of 0.881, while BLEU-4 decreased from 0.371 to 0.272. The decline in surface overlap coincided with stronger risk safety, stricter structure, and more guideline-directed outputs. Conclusions: Dynamic alignment shifted the model from linguistic mimicry toward clinically constrained HFrEF decision support. These findings suggest that staged optimization with policy alignment and retrieval grounding can improve evidence-based recommendations, while conventional language-overlap metrics may underestimate clinically safer generation.
科研通智能强力驱动
Strongly Powered by AbleSci AI