计算机科学
人工智能
判别式
特征提取
学习迁移
帧(网络)
模式识别(心理学)
特征(语言学)
卷积(计算机科学)
卷积神经网络
计算机视觉
混淆矩阵
数据集
人工神经网络
电信
哲学
语言学
作者
Kechuan Liu,Gaohao Zhou,Xin Feng Cheng
标识
DOI:10.1109/ipec54454.2022.9777334
摘要
Video content classification has a wide range of application scenarios. In situations such as public place security and behavior prediction, the type of behavior of a person is inferred from the continuous images in the video. Video data adds the concept of time compared to image data, which inevitably increases the overall computational effort. A typical research tool uses 3D convolution for feature map extraction in the time domain, but such methods exhibit significant performance loss due to frame down sampling. Inspired by the human judgment of video content, this study computes features on a single frame by transfer learning and then encodes the position of the resulting set of features. An attention mechanism is used to determine the keyframe locations. Finally, a multilayer perceptron is combined to achieve video content classification. The results of the study show that with the chosen dataset. Our model outperforms the 3D convolutional model in discriminative accuracy, confusion matrix, and down sampling conditions.
科研通智能强力驱动
Strongly Powered by AbleSci AI