计算机科学
注释
语义标注
人工智能
事件(粒子物理)
自然语言处理
情报检索
物理
量子力学
作者
Huaxiang Lu,Zhenlong Du
摘要
As a model of cross-media intelligence that combines computer vision and natural language processing, video semantic annotation facilitates automatic location of events in videos and describes video content in natural language. Unlike standard video annotation, densi video annotation requires the detection and description of multiple events in long videos, adding additional complexity to locate events in long videos. By proposing a dense video semantic annotation method based on deep learning, a single discrete tag sequence can be predicted for a given multimodal input, which includes a title tag for the event and a time tag representing the timestamp of the event. The proposed model uses unlabeled narrative videos for pre-training and uses transcribed speech and corresponding timestamps as a weakly supervised source of dense video annotation to replace manual annotation information, which expands the size of available datasets. Furthermore, by fine-tuning the model, we can apply it to the problem of paragraph annotation, generating paragraph descriptions about the entire video. The results show that the proposed model can predict high-quality event descriptions and relatively accurate time boundaries in different scenarios.
科研通智能强力驱动
Strongly Powered by AbleSci AI