计算机科学
插值(计算机图形学)
人工智能
图像缩放
计算机视觉
培训(气象学)
帧(网络)
图像(数学)
图像处理
计算机图形学(图像)
电信
物理
气象学
作者
Piotr Kopa Ostrowski,Daniel Węsierski,Anna Jezierska,Tomasz P. Stefański
标识
DOI:10.1109/tcsvt.2025.3575717
摘要
We introduce Frame Interpolation Pre-training (FIP), a simple learning technique for lifting deep image denoisers to video denoising with improved implicit temporal alignment. Modern video denoising networks typically rely on explicit motion estimation and alignment which are computationally intensive and harder to re-design and re-train, restricting their application scope and usability. Conversely, stacking frames and image denoisers, without incorporating explicit motion estimation modules, improves speed and benefits from a simpler design, thereby facilitating their generalizability to the video domain. However, it leads to lower accuracy due to suboptimal capture of temporal dependencies. To better leverage the adjacent frames in this setting and reduce the accuracy gap, we propose a novel training regime that divides the standard supervised training of the denoising task into two phases. In the initial phase, FIP guides the network to interpolate a fully masked central frame using only adjacent noisy input frames. In the subsequent phase, the pre-trained network is fine-tuned on denoising the central frame, now using all noisy input frames. Extensive diagnostics indicate that FIP-based networks provide better implicit motion estimation and temporal alignment. In effect, qualitative and quantitative evaluation on standard video denoising datasets with synthetic and real noise demonstrates that FIP consistently improves video denoising accuracy of motion-aware, video-lifted image denoisers without additional computational overhead during training and test time. Our code is available at https://github.com/camalab-ai/FIP.
科研通智能强力驱动
Strongly Powered by AbleSci AI