计算机科学
动作识别
块(置换群论)
编码(集合论)
人工智能
特征向量
帧速率
帧(网络)
变压器
模式识别(心理学)
电信
物理
电压
集合(抽象数据类型)
程序设计语言
量子力学
数学
班级(哲学)
几何学
作者
Gedas Bertasius,Heng Wang,Lorenzo Torresani
标识
DOI:10.48550/arxiv.2102.05095
摘要
We present a convolution-free approach to video classification built exclusively on self-attention over space and time. Our method, named "TimeSformer," adapts the standard Transformer architecture to video by enabling spatiotemporal feature learning directly from a sequence of frame-level patches. Our experimental study compares different self-attention schemes and suggests that "divided attention," where temporal attention and spatial attention are separately applied within each block, leads to the best video classification accuracy among the design choices considered. Despite the radically new design, TimeSformer achieves state-of-the-art results on several action recognition benchmarks, including the best reported accuracy on Kinetics-400 and Kinetics-600. Finally, compared to 3D convolutional networks, our model is faster to train, it can achieve dramatically higher test efficiency (at a small drop in accuracy), and it can also be applied to much longer video clips (over one minute long). Code and models are available at: https://github.com/facebookresearch/TimeSformer.
科研通智能强力驱动
Strongly Powered by AbleSci AI