计算机科学
人工智能
计算机视觉
语义学(计算机科学)
对象(语法)
分割
自然语言处理
特征(语言学)
图像分割
视频跟踪
视觉语言
可视化
背景(考古学)
作者
Ying Cao,Yu Wang,Lijuan Sun,Xiaomei Zou,Yuxiang Ma
标识
DOI:10.1016/j.ipm.2026.104712
摘要
Referring Video Object Segmentation (RVOS) aims to segment objects in videos based on language expressions. In this task, existing methods often face challenges due to semantic misalignment between language and visual modalities. To address this, this paper categorizes semantic misalignment into word-level, sentence-level, and cognitive-level issues. For these issues, we propose the Dynamic Alignment of Visual and Language Semantics (DAVLS) framework. In DAVLS, to address word-level semantic misalignment, we design the Word-level Semantic Dynamic Alignment (WSDA) to maximize the alignment of word-level semantics with appropriate visual features by embedding them into frame-level visual features and enabling cross-modal fusion visually. Based on accurate sentence-level object semantics generated by the textual adapter (TA), we propose the Sentence-level Semantic Dynamic Alignment (SSDA) to address sentence-level misalignment by aligning these semantics with the visual object state of the video clip. Due to the intrinsic linkage between WSDA and SSDA, they are integrated into a unified module, named SDA. Moreover, to mitigate the cognitive-level misalignment between humans and the segmentation model in understanding the semantics of referring objects, we introduce Semantic Augmentation (SA) to enhance the model’s comprehension of object semantics by expanding linguistic expressions. Simulation results on Ref-YouTube-VOS (3,978 videos), Ref-DAVIS17 (90 videos), A2D-Sentences (3,782 videos), and JHMDB-Sentences (928 videos) demonstrate the competitive performance of DAVLS, achieving a 0.9% improvement in J & F on Ref-YouTube-VOS and a 1.0% increase in Overall IoU on A2D-Sentences. Moreover, a statistically significant result ( p = . 03125 ), validated by the Wilcoxon signed-rank test, confirms the reliability of the performance gains.
科研通智能强力驱动
Strongly Powered by AbleSci AI