Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach

计算机科学 人工智能 变压器 语义学(计算机科学) 融合 编码(内存) 基本事实 一般化 机器学习 特征(语言学) 图像融合 人工神经网络 传感器融合 数据挖掘 特征提取 图像(数学) 模式识别(心理学) 特征选择 计算机视觉 概率逻辑 忠诚 融合机制 融合规则
作者
Jiayang Li,Chengjie Jiang,Junjun Jiang,Pengwei Liang,Jiayi Ma,Liqiang Nie
出处
期刊:IEEE Transactions on Pattern Analysis and Machine Intelligence [IEEE Computer Society]
卷期号:48 (4): 3970-3987 被引量:3
标识
DOI:10.1109/tpami.2025.3642842
摘要

Image fusion aims to blend complementary information from multiple sensing modalities, yet existing approaches remain limited in robustness, adaptability, and controllability. Most current fusion networks are tailored to specific tasks and lack the ability to flexibly incorporate user intent, especially in complex scenarios involving low-light degradation, color shifts, or exposure imbalance. Moreover, the absence of ground-truth fused images and the small scale of existing datasets make it difficult to train an end-to-end model that simultaneously understands high-level semantics and performs fine-grained multimodal alignment. We therefore present DiTFuse, an instruction-driven Diffusion Transformer (DiT) framework that performs end-to-end, semantics-aware fusion within a single model. By jointly encoding two images and natural-language instructions in a shared latent space, DiTFuse enables hierarchical and fine-grained control over fusion dynamics, overcoming the limitations of pre-fusion and post-fusion pipelines that struggle to inject high-level semantics. The training phase employs a multi-degradation masked-image modeling strategy, so the network jointly learns cross-modal alignment, modality-invariant restoration, and task-aware feature selection without relying on ground truth images. A curated, multi-granularity instruction dataset further equips the model with interactive fusion capabilities. DiTFuse unifies infrared-visible, multi-focus, and multi-exposure fusion-as well as text-controlled refinement and downstream tasks-within a single architecture. Experiments on public IVIF, MFF, and MEF benchmarks confirm superior quantitative and qualitative performance, sharper textures, and better semantic retention. The model also supports multi-level user control and zero-shot generalization to other multi-image fusion scenarios, including instruction-conditioned segmentation.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
bkagyin应助辛勤静珊采纳,获得10
刚刚
CipherSage应助快乐小兰采纳,获得10
刚刚
yeppp发布了新的文献求助10
刚刚
1秒前
戴路发布了新的文献求助10
2秒前
上官若男应助昆望采纳,获得10
2秒前
sw发布了新的文献求助10
2秒前
fxs完成签到,获得积分10
3秒前
4秒前
笑笑最可爱完成签到,获得积分10
4秒前
汉堡包应助llll采纳,获得10
4秒前
PWF完成签到,获得积分10
4秒前
高妍纯完成签到,获得积分10
5秒前
5秒前
spongxin完成签到,获得积分10
5秒前
xjl完成签到,获得积分10
6秒前
思源应助猫小咪采纳,获得10
7秒前
紫瓜完成签到,获得积分10
8秒前
贺鹏霖完成签到,获得积分10
8秒前
xjl发布了新的文献求助10
9秒前
9秒前
9秒前
星辰大海应助泠漓采纳,获得10
10秒前
10秒前
10秒前
bbq完成签到,获得积分10
11秒前
liujinjin完成签到,获得积分10
11秒前
taotie发布了新的文献求助10
11秒前
拾三完成签到,获得积分10
12秒前
吴咪发布了新的文献求助10
12秒前
13秒前
14秒前
辛勤静珊发布了新的文献求助10
15秒前
egret完成签到,获得积分10
15秒前
冬卿留完成签到,获得积分10
16秒前
小面包完成签到,获得积分10
16秒前
争气完成签到,获得积分10
16秒前
Una发布了新的文献求助10
16秒前
若一发布了新的文献求助50
16秒前
17秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Autoparametric Resonance in Mechanical Systems 1000
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 600
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 500
Auslegungsgeschichte 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7643039
求助须知:如何正确求助?哪些是违规求助? 9215999
关于积分的说明 19770731
捐赠科研通 7208275
什么是DOI,文献DOI怎么找? 3276465
关于科研通互助平台的介绍 2438211
邀请新用户注册赠送积分活动 2274313