计算机科学
安全性令牌
人工智能
一般化
突出
图像融合
背景(考古学)
代表(政治)
适配器(计算)
图像(数学)
计算机视觉
语义映射
模态(人机交互)
编码(集合论)
编码(内存)
融合
模式
融合规则
空间语境意识
语义鸿沟
解码方法
语义学(计算机科学)
联营
模式识别(心理学)
上下文模型
信息丢失
机器学习
医学影像学
作者
Chenyang Li,Rui Zhu,Hongyun Zhao,Xiongfei Li,Xi Zhang
标识
DOI:10.1016/j.patcog.2026.113848
摘要
Multimodal medical image fusion aims to combine complementary visual information from different imaging modalities to assist diagnosis and clinical decision-making. Existing methods often struggle to balance global semantic representation and local detail preservation, leading to blurred or incomplete features in salient regions. This work proposes ATDFusion (Adapter-Tuned Dual-Branch Network), featuring a global semantic branch and an auxiliary detail branch to jointly capture high-level context and fine-grained details. A Low-Rank Dynamic Token Adapter (LR-DTA) adaptively fine-tunes intermediate layers of pretrained models based on token number and rank. Additionally, a Fusion Enhancement Guidance (FEG) module imposes explicit spatial supervision via a saliency-aware loss to strictly preserve diagnostically critical regions. On three typical tasks (MRI-CT, MRI-PET, MRI-SPECT), ATDFusion surpasses state-of-the-art methods, improving SSIM and CC by 11.93% and 5.07%, respectively. Other metrics (e.g., EN, PSNR) also achieve leading results, validating its effectiveness. Furthermore, the model demonstrates strong zero-shot generalization on the unseen HECKTOR 2025 dataset. Code is available at https://github.com/pluto628/ATDFusion . • Asymmetric dual-branch framework resolves the semantic-detail trade-off. • Low-Rank Dynamic Token Adapter realizes efficient sample-specific fine-tuning. • Explicit spatial supervision guarantees the retention of metabolic/anatomical hotspots. • Achieves SOTA performance with strong zero-shot generalization on unseen datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI