计算机科学
行人
机器学习
跳跃式监视
人工智能
最小边界框
情态动词
特征(语言学)
传感器融合
数据建模
任务(项目管理)
数据挖掘
特征提取
运输工程
工程类
哲学
图像(数学)
数据库
经济
化学
管理
高分子化学
语言学
作者
Shengzhe Zhao,Haopeng Li,Qiuhong Ke,Liangchen Liu,Rui Zhang
标识
DOI:10.1109/lsp.2021.3134194
摘要
Pedestrian crossing intention prediction is crucial to traffic safety, which is a challenging task in real traffic scenarios. Traditional methods infer the intention of pedestrians to cross by predicting their future movements based on the observed trajectories in history. The performance of those methods is limited due to insufficient features and sources of information. To address those limitations, we propose a ViT-based model which incorporates multi-modal data to predict the pedestrian crossing intention. Specifically, the proposed model takes into consideration the visual information, poses, bounding box coordinates and action annotations, and gradually fuses those features for the final prediction. Besides, different data processing methods are designed based on the corresponding characteristics of different modalities to make full use of each type of data. Extensive ablation studies are conducted to show the performance of temporal modelling and feature fusion.
科研通智能强力驱动
Strongly Powered by AbleSci AI