计算机科学
编码器
人工智能
医学影像学
领域(数学分析)
钥匙(锁)
过程(计算)
自然语言处理
机器学习
课程
专家系统
医学知识
医学物理学
放射科
语义学(计算机科学)
领域知识
计算机视觉
医学
上下文图像分类
作者
Jiani Li,Zelan Li,Jianning Chi,Zhuming Bi
标识
DOI:10.1109/bibm66473.2025.11356033
摘要
Zero-shot medical image diagnosis is driven by vision-language pretraining on large-scale image-text pairs. However, aligning chest X-rays with reports remains challenging, particularly when multiple abnormalities with overlapping visual features lead to ambiguous correspondences. To address this challenge, we propose MedCMCL-VLP, a novel vision-language pretraining framework that simulates the hierarchical learning process of medical professionals. The framework consists of two key components: (1) a new enhanced Vision Mamba encoder with a spatial adaptor to capture long-range spatial dependencies and small-scale abnormalities in high-resolution chest X-rays and (2) a knowledge-guided text encoder that integrates medical domain knowledge with large language models (LLMs) to capture fine-grained semantic information from radiology reports. More importantly, MedCMCL-VLP employs a two-phase training paradigm that guides the model from novice to expert levels, effectively improving diagnostic accuracy. We evaluate MedCMCL-VLP on five public chest X-ray datasets. The experimental results show that our framework achieves state-of-the-art performance in both zero-shot and fine-tuned classification tasks.
科研通智能强力驱动
Strongly Powered by AbleSci AI