计算机科学
人工智能
分割
计算机视觉
特征(语言学)
光流
编码器
帧(网络)
架空(工程)
构造(python库)
卷积(计算机科学)
运动(物理)
模式识别(心理学)
图像(数学)
人工神经网络
操作系统
电信
哲学
程序设计语言
语言学
作者
Yifei Zhao,Xiaoying Wang,Junping Yin
标识
DOI:10.1109/jbhi.2025.3592897
摘要
Accurate and efficient Video Polyp Segmentation (VPS) is vital for the early detection of colorectal cancer and the effective treatment of polyps. However, achieving this remains highly challenging due to the inherent difficulty in modeling the spatial-temporal relationships within colonoscopy videos. Existing methods that directly associate video frames frequently fail to account for variations in polyp or background motion, leading to excessive noise and reduced segmentation accuracy. Conversely, approaches that rely on optical flow models to estimate motion and align frames incur significant computational overhead. To address these limitations, we propose a novel VPS framework, termed Deformable Alignment and Local Attention (DALA). In this framework, we first construct a shared encoder to jointly encode the feature representations of paired video frames. Subsequently, we introduce a Multi-Scale Frame Alignment (MSFA) module based on deformable convolution to estimate the motion between reference and anchor frames. The multi-scale architecture is designed to accommodate the scale variations of polyps arising from differing viewing angles and speeds during colonoscopy. Furthermore, Local Attention (LA) is employed to selectively aggregate the aligned features, yielding more precise spatial-temporal feature representations. Extensive experiments conducted on the challenging SUN-SEG dataset and PolypGen dataset demonstrate that DALA achieves superior performance compared to stateof-the-art models. The code will be publicly available at https://github.com/xff12138/DALA.
科研通智能强力驱动
Strongly Powered by AbleSci AI