自动汇总
计算机科学
块(置换群论)
人工智能
帧(网络)
弹丸
分割
利用
模式识别(心理学)
航程(航空)
机器学习
计算机视觉
化学
有机化学
电信
材料科学
几何学
数学
计算机安全
复合材料
作者
Wencheng Zhu,Jiwen Lu,Yucheng Han,Jie Zhou
标识
DOI:10.1016/j.patcog.2021.108312
摘要
In this paper, we propose a multiscale hierarchical attention approach for supervised video summarization. Different from most existing supervised methods which employ bidirectional long short-term memory networks, our method exploits the underlying hierarchical structure of video sequences and learns both the short-range and long-range temporal representations via a intra-block and a inter-block attention. Specifically, we first separate each video sequence into blocks of equal length and employ the intra-block and inter-block attention to learn local and global information, respectively. Then, we integrate the frame-level, block-level, and video-level representations for the frame-level importance score prediction. Next, we conduct shot segmentation and compute shot-level importance scores. Finally, we perform key shot selection to produce video summaries. Moreover, we extend our method into a two-stream framework, where appearance and motion information is leveraged. Experimental results on the SumMe and TVSum datasets validate the effectiveness of our method against state-of-the-art methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI