计算机科学
目标检测
人工智能
卷积神经网络
计算机视觉
任务(项目管理)
卷积(计算机科学)
对象(语法)
视频跟踪
Viola–Jones对象检测框架
跳跃式监视
分割
深度学习
对象类检测
领域(数学分析)
帧(网络)
模式识别(心理学)
人工神经网络
人脸检测
面部识别系统
数学分析
电信
数学
管理
经济
作者
Kai Kang,Wanli Ouyang,Hongsheng Li,Xiaogang Wang
摘要
Deep Convolution Neural Networks (CNNs) have shown impressive performance in various vision tasks such as image classification, object detection and semantic segmentation. For object detection, particularly in still images, the performance has been significantly increased last year thanks to powerful deep networks (e.g. GoogleNet) and detection frameworks (e.g. Regions with CNN features (RCNN)). The lately introduced ImageNet [6] task on object detection from video (VID) brings the object detection task into the video domain, in which objects' locations at each frame are required to be annotated with bounding boxes. In this work, we introduce a complete framework for the VID task based on still-image object detection and general object tracking. Their relations and contributions in the VID task are thoroughly studied and evaluated. In addition, a temporal convolution network is proposed to incorporate temporal information to regularize the detection results and shows its effectiveness for the task. Code is available at https://github.com/ myfavouritekk/vdetlib.
科研通智能强力驱动
Strongly Powered by AbleSci AI