Token-Mixer: Bind Image and Text in One Embedding Space for Medical Image Reporting

嵌入 图像(数学) 计算机科学 安全性令牌 空格(标点符号) 计算机视觉 人工智能 医学影像学 图像处理 计算机安全 操作系统
作者
Yan Yang,Jun Yu,Zhenqi Fu,Ke Zhang,Ting Yu,Xianyun Wang,Hanliang Jiang,Junhui Lv,Qingming Huang,Weidong Han
出处
期刊:IEEE Transactions on Medical Imaging [Institute of Electrical and Electronics Engineers]
卷期号:43 (11): 4017-4028 被引量:24
标识
DOI:10.1109/tmi.2024.3412402
摘要

Medical image reporting focused on automatically generating the diagnostic reports from medical images has garnered growing research attention. In this task, learning cross-modal alignment between images and reports is crucial. However, the exposure bias problem in autoregressive text generation poses a notable challenge, as the model is optimized by a word-level loss function using the teacher-forcing strategy. To this end, we propose a novel Token-Mixer framework that learns to bind image and text in one embedding space for medical image reporting. Concretely, Token-Mixer enhances the cross-modal alignment by matching image-to-text generation with text-to-text generation that suffers less from exposure bias. The framework contains an image encoder, a text encoder and a text decoder. In training, images and paired reports are first encoded into image tokens and text tokens, and these tokens are randomly mixed to form the mixed tokens. Then, the text decoder accepts image tokens, text tokens or mixed tokens as prompt tokens and conducts text generation for network optimization. Furthermore, we introduce a tailored text decoder and an alternative training strategy that well integrate with our Token-Mixer framework. Extensive experiments across three publicly available datasets demonstrate Token-Mixer successfully enhances the image-text alignment and thereby attains a state-of-the-art performance. Related codes are available at https://github.com/yangyan22/Token-Mixer.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
qawsedrf发布了新的文献求助10
刚刚
刚刚
曾经白安发布了新的文献求助10
刚刚
疯狂的白昼完成签到 ,获得积分10
1秒前
1秒前
molihuakai应助默默翩跹采纳,获得10
1秒前
欢喜板凳完成签到,获得积分10
1秒前
2秒前
今后应助Sunny采纳,获得10
2秒前
落后的静竹完成签到,获得积分10
2秒前
Akim应助西红柿采纳,获得10
2秒前
3秒前
3秒前
3秒前
imchenyin完成签到,获得积分10
3秒前
3秒前
wang发布了新的文献求助10
3秒前
3秒前
3秒前
粗鲁的男孩完成签到,获得积分10
3秒前
Nole应助kk采纳,获得30
3秒前
3秒前
Owen应助精明一寡采纳,获得10
3秒前
科研通AI6.4应助烂漫映之采纳,获得10
4秒前
Akim应助慈祥的丹寒采纳,获得10
4秒前
xiaobaiyang发布了新的文献求助10
4秒前
4秒前
bkagyin应助呵呵采纳,获得10
4秒前
5秒前
5秒前
5秒前
科研小白应助美女采纳,获得10
5秒前
JarryChao完成签到,获得积分10
5秒前
???发布了新的文献求助10
5秒前
5秒前
5秒前
空城发布了新的文献求助10
6秒前
6秒前
SciGPT应助英勇笑萍采纳,获得10
6秒前
6秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
The Multiple Self-States Drawing Technique 600
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Rosenblum, Global Change Biology 500
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7769181
求助须知:如何正确求助?哪些是违规求助? 9312341
关于积分的说明 20328283
捐赠科研通 7354496
什么是DOI,文献DOI怎么找? 3315979
关于科研通互助平台的介绍 2464892
邀请新用户注册赠送积分活动 2330552