代码本
图像复原
人工智能
模式识别(心理学)
特征(语言学)
计算机科学
代表(政治)
计算机视觉
特征提取
图像(数学)
编码(集合论)
图像处理
生成模型
数学
先验概率
迭代重建
过程(计算)
变压器
特征向量
面子(社会学概念)
源代码
隐马尔可夫模型
作者
Hongyu Li,Yu Liu,Tianyi Xu,Xiantong Zhen,Ran Gu,David Zhang,Jun Xu
标识
DOI:10.1109/tip.2026.3651985
摘要
Vector-Quantization (VQ) based discrete generative models are widely used to learn powerful high-quality (HQ) priors for blind image restoration (BIR). In this paper, we diagnose the side-effects of discrete VQ process essential to VQ-based BIR methods: 1) confining the representation capacity of HQ codebook, 2) being error-prone for code index prediction on low-quality (LQ) images, and 3) under-valuing the importance of input LQ image. These motivate us to learn continuous feature representation of HQ codebook for better restoration performance than using discrete VQ process. To further improve the restoration fidelity, we propose a new Self-in-Cross-Attention (SinCA) module to augment the HQ codebook with the feature of input LQ image, and perform cross-attention between LQ feature and input-augmented codebook. By this way, our SinCA leverages the input LQ image to enhance the representation of codebook for restoration fidelity. Experiments on four typical VQ-based BIR methods demonstrate that, by replacing the VQ process with a transformer using our SinCA, they achieve better quantitative and qualitative performance on blind image super-resolution and blind face restoration. The code and pre-trained models are publicly released at https://github.com/lhy-85/SinCA.
科研通智能强力驱动
Strongly Powered by AbleSci AI