计算机科学
计算机视觉
目标检测
人工智能
特征(语言学)
特征提取
对象(语法)
对象类检测
视频跟踪
图像分割
特征检测(计算机视觉)
雷达跟踪器
噪音(视频)
车辆跟踪系统
视觉对象识别的认知神经科学
模式识别(心理学)
图像处理
作者
Xu Cao,Yuqing Zhang,Huanxin Zou,Li Liu,Jun Li,Hao Chen,Xinyi Ying,Shitian He,Liyuan Pan
标识
DOI:10.1109/jstars.2026.3672667
摘要
In recent years, vehicle detection in unmanned aerial vehicle (UAV) videos has attracted significant attention. However, the traditional detection paradigm based on single frame image is faced with the performance bottleneck caused by rare pose, motion blur and occlusion. To address these issues, we propose a novel network comprising three important modules designed to fully exploit the spatiotemporal contextual information across video frames. Firstly, To strengthen the backbone network's capacity for extracting global contextual information from single-frame images, we propose the Visual State Space Context Module (VSSCM). By incorporating the 2D-Selective-Scan module (SS2D), VSSCM captures global dependencies and enriches contextual information without significantly increasing computational complexity. Secondly, the Temporal Information-Guided Spatial Attention Aggregation Module (TGSAM) is introduced to fuse features from critical regions in adjacent frames. Finally, the Self-Attention-based Classification Feature Aggregation Module (SACFAM) is designed to model the relation between object features across frames and perform feature aggregation based on the learned relation matrix, thereby effectively improving the quality of classification features in UAV videos. Extensive experiments are conducted on the challenging VisDrone2019-VID dataset, and the experimental results demonstrate the effectiveness and superiority of the proposed method.
科研通智能强力驱动
Strongly Powered by AbleSci AI