计算机科学
人工智能
稳健性(进化)
计算机视觉
判别式
分割
模式识别(心理学)
帧(网络)
上下文模型
背景(考古学)
图像分割
对象(语法)
电信
生物
化学
生物化学
古生物学
基因
作者
Jisheng Dang,Huicheng Zheng,Xiaohao Xu,Longguang Wang,Yulan Guo
标识
DOI:10.1109/tip.2024.3423390
摘要
Current video object segmentation approaches primarily rely on frame-wise appearance information to perform matching. Despite significant progress, reliable matching becomes challenging due to rapid changes of the object's appearance over time. Moreover, previous matching mechanisms suffer from redundant computation and noise interference as the number of accumulated frames increases. In this paper, we introduce a multi-frame spatio-temporal context memory (STCM) network to exploit discriminative spatio-temporal cues in multiple adjacent frames by utilizing a multi-frame context interaction module (MCI) for memory construction. Based on the proposed MCI module, a sparse group memory reader is developed to enable efficient sparse matching during memory reading. Our proposed method is generic and achieves state-of-the-art performance with real-time speed on benchmark datasets such as DAVIS and YouTube-VOS. In addition, our model exhibits robustness to sparse videos with low frame rates.
科研通智能强力驱动
Strongly Powered by AbleSci AI