计算机科学
心理学
自然语言处理
人工智能
认知心理学
模态(人机交互)
沟通
组分(热力学)
自然(考古学)
特征(语言学)
语言学
多模态
鉴定(生物学)
语义学(计算机科学)
代表(政治)
医学影像学
口译(哲学)
对比度(视觉)
作者
Tudor-Octavian Mihăiţă
标识
DOI:10.1109/synasc69064.2025.00074
摘要
Contrastive vision-language pretraining has shown strong performance in learning rich multimodal representations from image-text pairs. However, applying this paradigm to the medical domain presents significant challenges due to the complex visual patterns of medical conditions and the subtle, often ambiguous language used in radiology reports. To address these limitations, we propose CLIP-XRad, a weakly-supervised contrastive vision-language model specialized in chest X-ray interpretation. Our method introduces medical concept alignment into the contrastive learning process via a custom training objective that incorporates both instance-level and concept-level similarity, guiding the model to structure its embedding space around shared clinical semantics. Through extensive experiments, we demonstrate that CLIP-XRad achieves competitive performance in image-text retrieval, while also transferring effectively to downstream tasks such as multi-label classification and report generation. These results demonstrate that concept-aware contrastive pretraining can produce generalizable, semantically meaningful medical representations that provide a data-efficient solution for clinical applications.
科研通智能强力驱动
Strongly Powered by AbleSci AI