Text-Assisted Vision Model for Medical Image Segmentation

计算机视觉 计算机科学 图像分割 人工智能 分割 医学影像学 图像(数学) 计算机图形学(图像)
作者
Md. Motiur Rahman,Saeka Rahman,Smriti Bhatt,Miad Faezipour
出处
期刊:IEEE Journal of Biomedical and Health Informatics [Institute of Electrical and Electronics Engineers]
卷期号:29 (11): 8199-8212 被引量:4
标识
DOI:10.1109/jbhi.2025.3569491
摘要

Precise medical image segmentation is important for automating diagnosis and treatment planning in healthcare. While images present the most significant information for segmenting organs using deep learning models, text reports also provide complementary details that can be leveraged to improve segmentation precision. Performance improvement depends on the proper utilization of text reports and the corresponding images. Most attention modules focus on single-modality computation of spatial, channel, or pixel-level attention. They are ineffective in cross-modal alignment, raising issues in multi-modal scenarios. This study addresses these gaps by presenting a text-assisted vision (TAV) model for medical image segmentation with a novel attention computation module named tri-guided attention module (TGAM). TGAM computes visual-visual, language-language, and language-visual attention, enabling the model to understand the important features and correlation between images and medical notes. This module helps the model identify the relevant features within images, text annotations, and text annotations to visual interactions. We incorporate an attention gate (AG) that modulates the influence of TGAM, ensuring it does not overflow the encoded features with irrelevant or redundant information, while maintaining their uniqueness. We evaluated the performance of TAV on two popular datasets containing images and corresponding text annotations. We find TAV to be a new state-of-the-art model, as it improves the performance by 2-7% compared to other models. Extensive experiments were performed to demonstrate the effectiveness of each component of the proposed model.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
隐形曼青应助way_oz采纳,获得10
刚刚
必过六级完成签到,获得积分10
1秒前
呼噜噜ya完成签到 ,获得积分10
1秒前
子房发布了新的文献求助10
1秒前
英吉利25发布了新的文献求助10
2秒前
3秒前
3秒前
4秒前
小二郎应助坚强大神采纳,获得10
4秒前
薄饼哥丶发布了新的文献求助10
4秒前
5秒前
小鸭子发布了新的文献求助10
5秒前
5秒前
5秒前
goya完成签到,获得积分10
6秒前
6秒前
hqn完成签到,获得积分10
7秒前
7秒前
8秒前
11发布了新的文献求助10
8秒前
玉米泪先流完成签到,获得积分10
9秒前
9秒前
时嗷完成签到,获得积分10
9秒前
10秒前
DOC_XIONG应助zzh采纳,获得10
11秒前
Santino发布了新的文献求助10
11秒前
11秒前
hqn发布了新的文献求助10
11秒前
党蕊芳发布了新的文献求助10
11秒前
hjhhjh完成签到,获得积分10
12秒前
科研发布了新的文献求助10
12秒前
刻苦的千凝完成签到,获得积分20
12秒前
jgs发布了新的文献求助10
13秒前
星纪完成签到 ,获得积分10
13秒前
14秒前
傅礼貌完成签到,获得积分10
15秒前
15秒前
16秒前
四文鱼发布了新的文献求助10
16秒前
隐形曼青应助魁梧的人达采纳,获得10
18秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Principles of town planning: translating concepts to applications 1000
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The Great Hymn to Šamaš 500
Positive Obsession: The Life and Times of Octavia E. Butler 500
Interpolation and Regression Models for the Chemical Engineer: Solving Numerical Problems 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7692690
求助须知:如何正确求助?哪些是违规求助? 9253692
关于积分的说明 19985136
捐赠科研通 7265543
什么是DOI,文献DOI怎么找? 3291293
关于科研通互助平台的介绍 2447475
邀请新用户注册赠送积分活动 2296561