计算机科学
人工智能
算法
数据挖掘
模式识别(心理学)
噪音(视频)
算法设计
钥匙(锁)
数据建模
稳健性(进化)
信号处理
作者
Y. Chen,Huiqing Guo,Chunlei Peng,Nannan Wang,Xinbo Gao
标识
DOI:10.1109/tifs.2026.3678021
摘要
With the development of deep learning technology, the facial images generated by deepfake technology have reached a level of authenticity that is difficult to distinguish, posing a serious threat to personal privacy and data security. Therefore, it is of great significance to develop efficient and reliable deepfake detection technology. In recent years, Visual Language Models (VLM) have been applied to deepfake detection tasks due to their powerful multimodal understanding capabilities. However, the existing VLM have not been specifically optimized for deepfake detection tasks. When directly applied to this task, there are problems such as insufficient model accuracy and insufficient feature extraction, especially when dealing with complex forgery scenes. In response to these challenges, this paper proposes an innovative deepfake face detection method based on VLM and component-specific prompt tuning. We transform the deepfake detection task into a Visual Question Answering (VQA) task, making full use of the multimodal understanding capabilities of VLM and the flexibility of prompt tuning technology. This method uses a local prompt strategy to customize specific prompt questions for key facial components such as eyes, nose, and mouth, guiding the model to focus on the local features of these areas, thereby accurately capturing forgery traces. In addition, we introduced a feature extraction module Q-Former based on instructions, which can flexibly adjust the focus area of visual features according to prompts, significantly improving the model’s perception of locally forged features. By fusing these local features extracted by Q-Former and combining them with the language model to judge the authenticity of the overall face image, we can finally generate accurate prediction results. A large number of experimental results show that our method is significantly better than existing technologies in terms of detection accuracy and robustness.
科研通智能强力驱动
Strongly Powered by AbleSci AI