Surveillance Video-and-Language Understanding: From Small to Large Multimodal Models

计算机科学 人工智能 自然语言处理 多媒体 计算机视觉
作者
Tongtong Yuan,Xuange Zhang,Bo Liu,Kun Liu,Jian Gang Jin,Zhenzhen Jiao
出处
期刊:IEEE Transactions on Circuits and Systems for Video Technology [Institute of Electrical and Electronics Engineers]
卷期号:35 (1): 300-314 被引量:16
标识
DOI:10.1109/tcsvt.2024.3462433
摘要

Surveillance videos play a crucial role in public security. However, current tasks related to surveillance videos primarily focus on classifying and localizing anomalous events. Despite achieving notable performance, existing methods are restricted to detecting and classifying predefined events and lack satisfactory semantic understanding. To tackle this challenge, we introduce a novel research avenue focused on Video-and-Language Understanding for surveillance (VALU), and construct the first multimodal surveillance video dataset. We manually annotate the real-world surveillance dataset UCF-Crime with fine-grained event content and timing. Our newly annotated dataset, UCA (UCF-Crime Annotation), contains 23,542 sentences, with an average length of 20 words, and its annotated videos are as long as 110.7 hours. Moreover, we evaluate SOTA models on five multimodal tasks using this newly created dataset, establishing new baselines for surveillance VALU, from small to large models. Our experiments reveal that mainstream models, which perform well on previously public datasets, exhibit poor performance on surveillance video, highlighting new challenges in surveillance VALU. In addition to conducting baseline experiments to compare the performance of existing models, we also propose novel methods for multimodal anomaly detection tasks and finetune multimodal large language model models using our dataset. All the experiments highlight the necessity of constructing this multimodal dataset to advance surveillance AI. Upon the experimental results mentioned above, we conduct further in-depth analysis and discussion. The dataset and codes are provided athttps://xuange923.github.io/Surveillance-Video-Understanding.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
超帅冷雪发布了新的文献求助10
刚刚
韦恩发布了新的文献求助10
刚刚
Nowind发布了新的文献求助10
2秒前
星辰大海应助结实半邪采纳,获得10
2秒前
顾矜应助fjaa采纳,获得10
2秒前
3秒前
科研通AI6.3应助叶泠渊采纳,获得10
4秒前
5秒前
6秒前
韦恩完成签到,获得积分10
6秒前
活泼的橘子完成签到,获得积分10
10秒前
11秒前
完美世界应助kinder采纳,获得10
11秒前
123发布了新的文献求助10
12秒前
tiptip应助Eva采纳,获得30
12秒前
超帅冷雪完成签到,获得积分10
15秒前
15秒前
syyyy完成签到,获得积分10
15秒前
cxw陈祥薇发布了新的文献求助100
15秒前
matter完成签到 ,获得积分10
17秒前
18秒前
suye11111111111完成签到,获得积分10
18秒前
19秒前
19秒前
CipherSage应助Nena9采纳,获得10
21秒前
看文献了完成签到 ,获得积分10
21秒前
酷波er应助洁净的谷兰采纳,获得10
21秒前
21秒前
dudu发布了新的文献求助10
23秒前
Hello应助佳佳佳佳佳采纳,获得10
24秒前
24秒前
唯陌zero发布了新的文献求助10
26秒前
28秒前
齐天小圣完成签到 ,获得积分10
30秒前
KIKI完成签到,获得积分10
30秒前
Lucas应助雁子的女王陛下采纳,获得10
30秒前
31秒前
踏实一德应助唯一采纳,获得10
31秒前
不吃茄子完成签到 ,获得积分10
32秒前
lin完成签到,获得积分10
32秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Discerning Saints: Moralization of Intrinsic Motivation and Selective Prosociality at Work 500
Handbuch Trainingswissenschaft – Trainingslehre 500
Additive Manufacturing Design and Applications (ASM Handbook, Volume 24A) 500
Variations: A More Diverse Picture of Contemporary Art 400
Induction Heating and Heat Treatment (ASM Handbook, Volume 4C) 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7590396
求助须知:如何正确求助?哪些是违规求助? 9167820
关于积分的说明 19623142
捐赠科研通 7169551
什么是DOI,文献DOI怎么找? 3267307
关于科研通互助平台的介绍 2432173
邀请新用户注册赠送积分活动 2259518