Harnessing Text Insights With Visual Alignment for Medical Image Segmentation

图像分割 计算机视觉 计算机科学 人工智能 医学影像学 分割 图像(数学) 可视化
作者
Qingjie Zeng,Huan Luo,Zilin Lu,Yutong Xie,Zhiyong Wang,Yanning Zhang,Yong Xia
出处
期刊:IEEE Transactions on Medical Imaging [Institute of Electrical and Electronics Engineers]
卷期号:45 (2): 477-489 被引量:10
标识
DOI:10.1109/tmi.2025.3601359
摘要

Pre-trained vision-language models (VLMs) and language models (LMs) have recently garnered significant attention due to their remarkable ability to represent textual concepts, opening up new avenues in vision tasks. In medical image segmentation, efforts are being made to integrate text and image data using VLMs and LMs. However, current text-enhanced approaches face several challenges. First, using separate pre-trained vision and text models to encode image and text data can result in semantic shifts. Second, while VLMs can establish the correspondence between visual and textual features when pre-trained on paired image-text data, this alignment often deteriorates during segmentation tasks due to misalignment between the text and vision components in ongoing learning. In this paper, we propose TeViA, a novel approach that seamlessly integrates with various vision and text models, irrespective of their pre-training relationships. This integration is achieved through a segmentation-specific text-to-vision alignment design, ensuring both information gain and semantic consistency. Specifically, for each training data, a foreground visual representation is extracted from the segmentation head and used to supervise projection layers, thereby adjusting the textual features to better contribute to the segmentation task. Additionally, a historic visual prototype is created by aggregating target semantics from all training data and is updated using a momentum-based manner. This prototype aims to enhance the visual representation of each data instance by establishing feature-level connections, which in turn refines the textual features. The superiority of TeViA is validated on five public datasets, exhibiting over 6% Dice improvements compared to vision-only methods. Code is available at: https://github.com/jgfiuuuu/TeViA.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
无敌的齐天大圣完成签到 ,获得积分10
刚刚
看不懂文献咕咕嘎嘎完成签到,获得积分10
2秒前
2秒前
3秒前
热闹的冬天完成签到,获得积分10
3秒前
5秒前
6秒前
Akim应助刘骁萱采纳,获得10
6秒前
小目标发布了新的文献求助10
7秒前
8秒前
俏皮的老城完成签到 ,获得积分0
9秒前
渡春屿完成签到 ,获得积分10
9秒前
今后应助Liranran采纳,获得10
10秒前
Rae完成签到 ,获得积分10
10秒前
体贴鱼完成签到,获得积分10
10秒前
黄景滨完成签到 ,获得积分10
10秒前
Doctor发布了新的文献求助10
11秒前
11秒前
备用完成签到 ,获得积分10
11秒前
天天快乐应助棋士采纳,获得10
12秒前
12秒前
12秒前
Yuh完成签到 ,获得积分10
13秒前
科研通AI2S应助小目标采纳,获得10
13秒前
Linly发布了新的文献求助80
14秒前
阿司匹林完成签到,获得积分10
14秒前
15秒前
15秒前
ALBERT发布了新的文献求助10
15秒前
16秒前
慕青应助渴望者采纳,获得10
16秒前
研友_nPbeR8完成签到,获得积分10
17秒前
17秒前
虚拟的涟妖完成签到 ,获得积分10
19秒前
bruce完成签到,获得积分10
19秒前
Hansen完成签到,获得积分10
20秒前
刘骁萱发布了新的文献求助10
20秒前
21秒前
棋士发布了新的文献求助10
23秒前
春风得意发布了新的文献求助10
23秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Health Psychology 800
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Electric machines: theory, operating applications, and controls 500
The Analytical and Numerical Solution of Electric and Magnetic Fields 500
When Is Two-Stage Sample Robust Optimization Asymptotically Optimal? 500
Discerning Saints: Moralization of Intrinsic Motivation and Selective Prosociality at Work 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7593471
求助须知:如何正确求助?哪些是违规求助? 9170643
关于积分的说明 19629353
捐赠科研通 7171301
什么是DOI,文献DOI怎么找? 3267609
关于科研通互助平台的介绍 2432450
邀请新用户注册赠送积分活动 2260262