高保真
计算机科学
忠诚
人机交互
多媒体
计算机图形学(图像)
工程类
电信
电气工程
作者
Chao Xu,Yang Liu,Jiazheng Xing,Weida Wang,Mingze Sun,Jun Dan,Tianxin Huang,Siyuan Li,Zhi-Qi Cheng,Ying Yu Tai,Baigui Sun
出处
期刊:Proceedings
[Institute of Electrical and Electronics Engineers]
日期:2024-06-16
卷期号:: 1292-1302
被引量:15
标识
DOI:10.1109/cvpr52733.2024.00129
摘要
In this paper, we abstract the process of people hearing speech, extracting meaningful cues, and creating vari-ous dynamically audio-consistent talking faces, termed Lis-tening and Imagining, into the task of high-fidelity diverse talking faces generation from a single audio. Specifically, it involves two critical challenges: one is to effectively de-couple identity, content, and emotion from entangled au-dio, and the other is to maintain intra-video diversity and inter- video consistency. To tackle the issues, we first dig out the intricate relationships among facial factors and sim-plify the decoupling process, tailoring a Progressive Audio Disentanglement for accurate facial geometry and seman-tics learning, where each stage incorporates a customized training module responsible for a specific factor. Secondly, to achieve visually diverse and audio-synchronized animation solely from input audio within a single model, we intro-duce the Controllable Coherent Frame generation, which involves the flexible integration of three trainable adapters with frozen Latent Diffusion Models (LDMs) to focus on maintaining facial geometry and semantics, as well as texsture and temporal coherence between frames. In this way, we inherit high-quality diverse generation from LDMs while significantly improving their controllability at a low training cost. Extensive experiments demonstrate the flexibility and effectiveness of our method in handling this paradigm. The codes will be released at FaceChain.
科研通智能强力驱动
Strongly Powered by AbleSci AI