人工智能
计算机科学
图像缩放
插值(计算机图形学)
计算机视觉
特征(语言学)
卷积神经网络
脱模
特征提取
深度学习
模式识别(心理学)
图像处理
图像(数学)
彩色图像
语言学
哲学
作者
Hui Liu,Gongguan Chen,Meng Liu,Liqiang Nie
标识
DOI:10.1109/tcsvt.2024.3409395
摘要
Image sequence interpolation is a critical research area in computer vision with broad applications in video frame interpolation and medical image interlayer interpolation. Traditional deep learning-based methods in this domain predominantly rely on deep convolutional neural networks (CNNs), which, despite their effectiveness, are limited by the inherent constraints of CNN architecture, impacting their interpolation accuracy. To address these limitations, we introduce the Pre-ISIformer, a parallel multi-channel adaptive image sequence interpolation network founded on pre-trained transformers. This innovative network is composed of three integral modules: 1) Global feature extraction module is designed to extract primary features from the input images using a pre-trained Swin-transformer model, ensuring comprehensive global feature coverage. 2) Feature sequence construction module adaptively decomposes the object’s motion path across different frames, facilitating a detailed analysis of motion dynamics. And 3) Intermediate image reconstruction module is responsible for accurately capturing target displacements. Furthermore, we incorporate distinct metrics for pixel loss and gradient loss to meticulously reconstruct the texture and contours of the intermediate images. Our network has been rigorously tested on various datasets for two primary applications: video frame interpolation and interlayer interpolation in medical imaging. The results from these experiments showcase the superior performance and effectiveness of the Pre-ISIformer, establishing it as a significant advancement in the field of image sequence interpolation.
科研通智能强力驱动
Strongly Powered by AbleSci AI