计算机科学
判别式
语音识别
人工智能
韵律
持续时间(音乐)
自然语言处理
预处理器
特征(语言学)
自然(考古学)
发音
语言模型
自然语言
钥匙(锁)
计算机视觉
隐藏字幕
领域(数学分析)
机器翻译
地标
特征提取
眼动
基本事实
任务分析
作者
Z L Zhang,Liang Li,Gaoxiang Cong,Chunshan Liu,Yuhan Gao,Xiaowan Wang,Tao Gu,Yuankai Qi
标识
DOI:10.48550/arxiv.2512.17154
摘要
Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's visual performance. However, existing alignment approaches based on visual features face two key limitations: (1)they rely on complex, handcrafted visual preprocessing pipelines, including facial landmark detection and feature extraction; and (2) they generalize poorly to unseen visual domains, often resulting in degraded alignment and dubbing quality. To address these issues, we propose InstructDubber, a novel instruction-based alignment dubbing method for both robust in-domain and zero-shot movie dubbing. Specifically, we first feed the video, script, and corresponding prompts into a multimodal large language model to generate natural language dubbing instructions regarding the speaking rate and emotion state depicted in the video, which is robust to visual domain variations. Second, we design an instructed duration distilling module to mine discriminative duration cues from speaking rate instructions to predict lip-aligned phoneme-level pronunciation duration. Third, for emotion-prosody alignment, we devise an instructed emotion calibrating module, which finetunes an LLM-based instruction analyzer using ground truth dubbing emotion as supervision and predicts prosody based on the calibrated emotion analysis. Finally, the predicted duration and prosody, together with the script, are fed into the audio decoder to generate video-aligned dubbing. Extensive experiments on three major benchmarks demonstrate that InstructDubber outperforms state-of-the-art approaches across both in-domain and zero-shot scenarios.
科研通智能强力驱动
Strongly Powered by AbleSci AI