计算机科学
文本生成
水准点(测量)
面子(社会学概念)
情报检索
相关性(法律)
自然语言处理
质量(理念)
比例(比率)
语料库
人工智能
语言学
哲学
物理
法学
认识论
地理
量子力学
政治学
大地测量学
作者
Jianhui Yu,Hao Zhu,Liming Jiang,Chen Change Loy,Weidong Cai,Wayne Wu
出处
期刊:
日期:2023-06-01
被引量:14
标识
DOI:10.1109/cvpr52729.2023.01422
摘要
Text-driven generation models are flourishing in video generation and editing. However, face-centric text-to-video generation remains a challenge due to the lack of a suitable dataset containing high-quality videos and highly relevant texts. This paper presents Celeb V- Text, a large-scale, di-verse, and high-quality dataset of facial text-video pairs, to facilitate research on facial text-to- video generation tasks. CelebV-Text comprises 70,000 in-the-wild face video clips with diverse visual content, each paired with 20 texts gen-erated using the proposed semi-automatic text generation strategy. The provided texts are of high quality, describing both static and dynamic attributes precisely. The supe-riority of CelebV- Text over other datasets is demonstrated via comprehensive statistical analysis of the videos, texts, and text-video relevance. The effectiveness and potential of CelebV- Text are further shown through extensive self-evaluation. A benchmark is constructed with representative methods to standardize the evaluation of the facial text-to-video generation task. All data and models are publicly available 1 1 Project page: https://celebv-text.github.io.
科研通智能强力驱动
Strongly Powered by AbleSci AI