计算机科学
计算机视觉
人工智能
变压器
图像复原
图像(数学)
计算机图形学(图像)
图像处理
工程类
电气工程
电压
作者
Monika Kwiatkowski,Simon Matern,Olaf Hellwich
出处
期刊:
日期:2025-02-26
卷期号:: 1383-1391
标识
DOI:10.1109/wacv61041.2025.00142
摘要
Most deep-learning models for vision tasks rely on RGB images as their primary input layer, assuming the model inherently discovers an optimal representation. In this work, we challenge this assumption and show that image gradients offer a straightforward yet robust representation for multi-frame image restoration. We demonstrate that clusters naturally emerge within gradient patches, indicating improved estimation of the underlying signal. We develop a Video Swin-Transformer model operating in the gradient domain, facilitated by the implementation of two differentiable gradient modules. One module computes image gradients using convolutions with gradient filters, while the other reconstructs an RGB image from its gradient representation using deconvolution in the frequency domain. Additionally, we employ a composite training loss that measures the error both in the color domain and its gradient counterpart. Applied to a multi-frame image restoration task involving the removal of lighting, shadows, and occlusions, our model consistently outperforms RGB-based counterparts without introducing additional parameters, thanks to its gradient regularization. We further apply our framework to various restoration tasks, discussing its advantages and limitations. Qualitative results highlight the model's improved generalization to real-world video scenarios, demonstrating successful adaptation from synthetic image training to real video data deployment.
科研通智能强力驱动
Strongly Powered by AbleSci AI