对抗制
一致性(知识库)
发电机(电路理论)
身份(音乐)
运动(物理)
计算机科学
主题(文档)
计算机视觉
生成语法
人工智能
透视图(图形)
面部识别系统
人机交互
面子(社会学概念)
人工神经网络
面部表情
GSM演进的增强数据速率
模式识别(心理学)
生成对抗网络
运动捕捉
会话(web分析)
作者
Antonio Greco,Nicola Strisciuglio,Mario Vento
标识
DOI:10.1109/taffc.2025.3634204
摘要
Artificially applying specific emotions to videos of people faces with a neutral expression, while preserving the identity of the subject is a challenging task. When parts of the face are synthetically moved to generate an emotion, it typically results in spatio-temporal artifacts in the generated videos, or inconsistency to preserve the identity of subjects. Existing methods that deploy spatio-temporal convolutions and de-convolutions to generate consecutive frames in a single step are not able to ensure proper motion dynamics, in the sense that the emotion may be not visible on the face or the facial features are distorted in the video. At the same time, approaches that generate motion and identity in two separate steps are not able to ensure the consistency of the subject identity after the generation of the emotion. In this paper we propose a novel method, Video Identity-Consistent Emotion GAN (VICEGAN), that improves the video generative capabilities of two-step methods. We decouple motion and content generation, thus ensuring the consistency of subject identity in the generated videos by using an encoder-decoder generator and a new identity-preserving loss in an adversarial framework. The proposed neural network architecture also guarantees the generation of proper motion of the target expressions, mitigating the presence of artifacts. We evaluated VICEGAN on the MUG dataset and compared it with a method based on a GAN, ImaGINator, demonstrating superior performance both quantitatively and qualitatively, and with a popular method based on a diffusion model, LFDM, showing a better capability to generate recognizable emotions.
科研通智能强力驱动
Strongly Powered by AbleSci AI