计算机科学
变压器
遥感
红外线的
人工智能
计算机视觉
翻译(生物学)
地质学
电气工程
电压
光学
工程类
物理
信使核糖核酸
基因
生物化学
化学
作者
Zonghao Han,Xiaoning Chen,Zixiang Ye,Yuru Su,Lefan Wang,Shaohui Mei
标识
DOI:10.1109/tgrs.2025.3604474
摘要
Aerial visible-to-infrared image translation technology aims to generate infrared images from visible inputs, effectively expanding the acquisition of infrared images in complex aerial remote sensing scenarios. Although existing CNN-based approaches demonstrate proficiency in capturing local spatial details, they often fail to model global scene context, which is crucial for maintaining structural consistency in complex aerial scenes. To address this challenge, a hybrid architecture network is proposed to model long-range dependencies while preserving fine-grained local details via an encoder-decoder framework. We further introduce a Dynamic Parallel Window-based Attention mechanism, which dynamically parallelizes window-based and shifted window-based multi-head self-attention in separate streams across transformer blocks to enhance global context modeling. Additionally, a masked image pre-training framework with wavelet transform loss is designed to guide multi-scale feature alignment and high-frequency detail reconstruction, effectively addressing the texture discrepancy between visible and infrared modalities. Extensive experiments conducted on the upgraded AVIID and DroneVehicle benchmark datasets demonstrate that our method significantly outperforms current state-of-the-art approaches in terms of both visual quality and perceptual metrics. The code of USTNet and the upgraded AVIID dataset are available at https://github.com/silver-hzh/USTNet.
科研通智能强力驱动
Strongly Powered by AbleSci AI