Enhancing medical vision-language contrastive learning via inter-matching relation modelling

关系（数据库）计算机科学匹配（统计）人工智能计算机视觉自然语言处理医学影像学数学数据挖掘统计

作者

Mingjian Li,Mingyuan Meng,Michael Fulham,Dagan Feng,Lei Bi,Jinman Kim

出处

期刊：IEEE Transactions on Medical Imaging [Institute of Electrical and Electronics Engineers]
日期：2025-01-01 卷期号：: 1-1 被引量：2

链接

nih.govdoi.org

标识

DOI：10.1109/tmi.2025.3534436

摘要

Medical image representations can be learned through medical vision-language contrastive learning (mVLCL) where medical imaging reports are used as weak supervision through image-text alignment. These learned image representations can be transferred to and benefit various downstream medical vision tasks such as disease classification and segmentation. Recent mVLCL methods attempt to align image sub-regions and the report keywords as local-matchings. However, these methods aggregate all local-matchings via simple pooling operations while ignoring the inherent relations between them. These methods therefore fail to reason between local-matchings that are semantically related, e.g., local-matchings that correspond to the disease word and the location word (semantic-relations), and also fail to differentiate such clinically important local-matchings from others that correspond to less meaningful words, e.g., conjunction words (importance-relations). Hence, we propose a mVLCL method that models the inter-matching relations between local-matchings via a relation-enhanced contrastive learning framework (RECLF). In RECLF, we introduce a semantic-relation reasoning module (SRM) and an importance-relation reasoning module (IRM) to enable more fine-grained report supervision for image representation learning. We evaluated our method using six public benchmark datasets on four downstream tasks, including segmentation, zero-shot classification, linear classification, and cross-modal retrieval. Our results demonstrated the superiority of our RECLF over the state-of-the-art mVLCL methods with consistent improvements across single-modal and cross-modal tasks. These results suggest that our RECLF, by modelling the inter-matching relations, can learn improved medical image representations with better generalization capabilities.

求助该文献

最长约 10秒，即可获得该文献文件

Enhancing medical vision-language contrastive learning via inter-matching relation modelling

今日热心研友