计算机科学
一致性(知识库)
一般化
光学(聚焦)
蒸馏
人工智能
机器学习
数据挖掘
身份(音乐)
先验与后验
情报检索
利用
自然语言处理
模式识别(心理学)
建筑
钥匙(锁)
内容(测量理论)
度量(数据仓库)
语义学(计算机科学)
作者
Beijia Sun,Haiqing Du,Yiwei Fu,Tingya Dong,Wenzhe Lu
标识
DOI:10.1109/vcip67698.2025.11396833
摘要
The advent of deepfake technology enables the synthesis of highly realistic audio-visual content, posing severe challenges such as identity impersonation and public opinion manipulation. Existing detection methods primarily focus on unimodal analysis or shallow consistency modeling, making it difficult to effectively address cross-modal forgeries and well-synchronized fake samples. These limitations result in poor generalization and high misclassification rates. To tackle these issues, this paper proposes an innovative multi-teacher knowledge distillation detection framework based on content consistency and semantic consistency. By leveraging high-level supervision from speech recognition and semantic comprehension, the framework guides the detection model to simultaneously learn cross-modal alignment in terms of both content and semantics. Additionally, a Transformer-based architecture is integrated to enhance the audio temporal modeling capability of the detection model, improving its ability to capture long-term temporal dependencies in deepfake speech. Comprehensive evaluations conducted on the FakeAVCeleb dataset demonstrate that the proposed method outperforms existing methods in terms of accuracy and AUC, while also exhibiting fewer parameters with higher computational efficiency, highlighting strong practical applicability.
科研通智能强力驱动
Strongly Powered by AbleSci AI