计算机科学
变压器
图像处理
电子工程
图像(数学)
计算机视觉
人工智能
信号处理
电气工程
特征提取
图像分割
电压
图像压缩
算法设计
作者
Xue Wu,Kaibing Zhang,Jingwei Xin,Nannan Wang,Xinbo Gao
标识
DOI:10.1109/tmm.2026.3680070
摘要
Vision Transformer-based models have shown promising performance on various image recovery tasks including image super-resolution (SR). However, their intensive computational complexity hinders the applicability, particularly in resource-limited scenarios. To alleviate this weakness, developing a lightweight yet effective transformer model is worthy of further investigation. In this paper, we elaborate on a novel Window-Free Transformer (WFT) for efficient SR. The WFT consists of three core components, namely multi-head multi-scale spatial self-attention (MMSSA), multi-head channel self-attention (MCSA), and dual spatial-gate feed-forward network (DSGFN). Specifically, the MMSSA achieves long-range dependencies by splitting the attention head into different groups and each group performs spatial self-attention across different scales. The MCSA performs global channel self-attention by measuring attention scores between different channel tokens to strengthen feature interactions across all channels. MMSSA and MCSA can conquer the bottleneck of window-constraint schemes and yield sufficient receptive fields. Additionally, the DSGFN is designed to enhance local contextual information via a dual spatial-gate mechanism. Thorough experimental evaluations demonstrate that the WFT outperforms the existing leading SR methods on five benchmark datasets in both quantitative and qualitative assessments while maintaining a compact model size. Specifically, WFT outperforms SwinIR-light and SRFormer-light by up to 0.38dB and 0.33dB in PSNR on the Urban100 dataset, and by up to 0.41dB and 0.39dB on the Manga109 dataset, respectively, while requiring fewer parameters and computations. Moreover, our lightweight WFT can be scaled to the classical SR task, delivering remarkable SR performance over other competitors. To facilitate reproducibility, we will release the source code upon publication.
科研通智能强力驱动
Strongly Powered by AbleSci AI