Softmax函数
计算机科学
保险丝(电气)
卷积神经网络
联营
人工智能
卷积(计算机科学)
动作识别
图层(电子)
模式识别(心理学)
卷积码
网络体系结构
班级(哲学)
人工神经网络
解码方法
电信
电气工程
工程类
计算机安全
有机化学
化学
作者
Christoph Feichtenhofer,Axel Pinz,Andrew Zisserman
标识
DOI:10.1109/cvpr.2016.213
摘要
Recent applications of Convolutional Neural Networks (ConvNets) for human action recognition in videos have proposed different solutions for incorporating the appearance and motion information. We study a number of ways of fusing ConvNet towers both spatially and temporally in order to best take advantage of this spatio-temporal information. We make the following findings: (i) that rather than fusing at the softmax layer, a spatial and temporal network can be fused at a convolution layer without loss of performance, but with a substantial saving in parameters, (ii) that it is better to fuse such networks spatially at the last convolutional layer than earlier, and that additionally fusing at the class prediction layer can boost accuracy, finally (iii) that pooling of abstract convolutional features over spatiotemporal neighbourhoods further boosts performance. Based on these studies we propose a new ConvNet architecture for spatiotemporal fusion of video snippets, and evaluate its performance on standard benchmarks where this architecture achieves state-of-the-art results.
科研通智能强力驱动
Strongly Powered by AbleSci AI