人工智能
计算机科学
计算机视觉
图像(数学)
模式识别(心理学)
图像处理
图像分割
特征提取
迭代重建
像素
上下文图像分类
目标检测
图像配准
特征(语言学)
图像复原
医学影像学
作者
Yupeng Zhou,Zhen Li,Chun-Le Guo,Li Liu,Ming‐Ming Cheng,Qibin Hou
标识
DOI:10.1109/tpami.2026.3685679
摘要
Previous works have shown that increasing the window size for Transformer-based image super-resolution models (e.g., SwinIR) can significantly improve the model performance. Still, the computation overhead is also considerable when the window size gradually increases. In this paper, we present SRFormer, a simple but novel method that can enjoy the benefit of large window self-attention but introduces even less computational burden. The core of our SRFormer is the permuted self-attention (PSA), which strikes an appropriate balance between the channel and spatial information for self-attention. Without any bells and whistles, we show that our SRFormer achieves a 33.86dB PSNR score on the Urban100 dataset, which is 0.46dB higher than that of SwinIR but uses fewer parameters and computations. In addition, we also attempt to scale up the model by further enlarging the window size and channel numbers to explore the potential of Transformer-based models. Experiments show that our scaled model, named SRFormerV2, can further improve the results and achieves state-of-the-art. We hope our simple and effective approach could be useful for future research in super-resolution model design. The homepage is https://z-yupeng.github.io/SRFormer/.
科研通智能强力驱动
Strongly Powered by AbleSci AI