清晨好,您是今天最早来到科研通的研友!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您科研之路漫漫前行!

Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding

计算机科学 人工智能 计算机视觉 自然语言处理 多媒体
作者
Xiao Wang,Jianlong Wu,Zijia Lin,Fuzheng Zhang,Di Zhang,Liqiang Nie
出处
期刊:IEEE Transactions on Pattern Analysis and Machine Intelligence [IEEE Computer Society]
卷期号:47 (4): 2912-2923 被引量:1
标识
DOI:10.1109/tpami.2025.3528394
摘要

Recently, video-language understanding has achieved great success through large-scale pre-training. However, data scarcity remains a prevailing challenge. This study quantitatively reveals an "impossible trinity" among data quantity, diversity, and quality in pre-training datasets. Recent efforts seek to refine large-scale, diverse ASR datasets compromised by low quality through synthetic annotations. These methods successfully refine the original annotations by leveraging useful information in multimodal video content (frames, tags, ASR transcripts, etc.). Nevertheless, they struggle to mitigate noise within synthetic annotations and lack scalability as the dataset size expands. To address these issues, we introduce the Video DataFlywheel framework, which iteratively refines video annotations with improved noise control methods. For iterative refinement, we first leverage a video-language model to generate synthetic annotations, resulting in a refined dataset. Then, we pre-train on it and fine-tune on human refinement examples for a stronger model. These processes are repeated for continuous improvement. For noise control, we present AdaTaiLr, a novel method that requires weaker assumptions on noise distribution. This method proves more effective in large datasets and offers theoretical guarantees. The combination of iterative refinement and AdaTaiLr can achieve better scalability in video-language understanding. Extensive experiments show that our framework outperforms existing data refinement baselines, delivering a 3% performance boost and improving dataset quality with minimal diversity loss. Furthermore, our refined dataset facilitates significant improvements in various video-language understanding tasks, including video question answering and text-video retrieval.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
3秒前
19秒前
anugraphics发布了新的文献求助40
24秒前
陶醉妙松完成签到,获得积分10
29秒前
30秒前
anugraphics发布了新的文献求助40
34秒前
47秒前
悲凉的丝完成签到,获得积分10
58秒前
59秒前
anugraphics发布了新的文献求助40
1分钟前
anugraphics发布了新的文献求助40
1分钟前
可爱花瓣发布了新的文献求助10
1分钟前
景妙海完成签到 ,获得积分10
1分钟前
anugraphics发布了新的文献求助40
1分钟前
愉快初曼完成签到,获得积分10
1分钟前
Benhnhk21完成签到,获得积分10
1分钟前
1分钟前
柚子茶应助369ninja采纳,获得10
1分钟前
1分钟前
1分钟前
神勇寄风完成签到,获得积分10
2分钟前
土豆泥完成签到,获得积分10
2分钟前
77发布了新的文献求助10
2分钟前
个性的一手完成签到,获得积分10
2分钟前
大个应助77采纳,获得50
2分钟前
柚子茶应助369ninja采纳,获得10
2分钟前
火星上安柏完成签到,获得积分10
2分钟前
糟糕的问丝完成签到,获得积分10
2分钟前
2分钟前
2分钟前
俭朴映寒完成签到,获得积分10
3分钟前
柚子茶应助369ninja采纳,获得10
3分钟前
可爱花瓣发布了新的文献求助10
3分钟前
整齐诺言完成签到,获得积分10
3分钟前
3分钟前
小螃蟹完成签到 ,获得积分10
3分钟前
合适的飞绿完成签到,获得积分10
3分钟前
明理傥完成签到,获得积分10
3分钟前
king完成签到 ,获得积分10
3分钟前
柚子茶应助369ninja采纳,获得10
4分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Principles of town planning: translating concepts to applications 1000
Navigating Normative Orders. Interdisciplinary Perspectives 800
1 Peter and Christ's Descent to the Dead in Its Early Christian Reception 700
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7738853
求助须知:如何正确求助?哪些是违规求助? 9287793
关于积分的说明 20184844
捐赠科研通 7316768
什么是DOI,文献DOI怎么找? 3306016
关于科研通互助平台的介绍 2458383
邀请新用户注册赠送积分活动 2315935