单眼
计算机科学
计算机视觉
人工智能
跟踪(教育)
心理学
教育学
作者
Hui Xiong,Fang Zhang,Xuzhi Fang,Heye Huang,Quan Yuan,Qinggui Pan,Dezhao Zhang
标识
DOI:10.1109/itsc58415.2024.10920028
摘要
In the autonomous driving environment, Vulnerable Road Users (VRUs), including pedestrians and various types of riders with diverse visual attributes and unpredictable movements, necessitate prompt and precise motion intention to ensure their safety. However, detection methods tend to misidentify VRUs, while tracking methods struggle with maintaining fast-moving objects, and prediction methods overly depend on map topological data, which leads to unsatisfactory accuracy and reliability in VRU prediction. This paper introduces a DTP-M3Net for VRU-oriented monocular multiclass multistage detection, tracking and prediction. DTP-M3Net makes use of the basic location and classification information, as well as historical trajectories and coarse-to-fine semantic features to achieve comprehensive perception. Key innovations encompass a faster multi-object detection via deep neural network, an online multi-object tracking with ego-motion compensation, and a joint trajectory prediction based on a sequence-to-sequence network and spatial-temporal predicted cues. Excellently, deep convolutional features generated by detection are shared among downstream tracking and prediction modules. The effectiveness of proposed DTP-M3Net is validated on the public MOT challenge and VRU-Track dataset, demonstrating improvements in VRU perception.
科研通智能强力驱动
Strongly Powered by AbleSci AI