计算机科学
稳健性(进化)
异步通信
人工智能
深度学习
生成语法
特征学习
同步(交流)
特征提取
利用
建筑
机器学习
特征(语言学)
隐马尔可夫模型
数字签名
数字取证
模糊测试
语音识别
电子学习
生成模型
隐藏字幕
图层(电子)
封面(代数)
变压器
人机交互
作者
Truong Xuan Hung,Luong The Dung,Tran Anh Tu
出处
期刊:Nghiên cứu khoa học và công nghệ trong lĩnh vực an toàn thông tin
[Information Security Journal]
日期:2026-08-16
卷期号:: 5-26
标识
DOI:10.54654/isj.v2i28.1248
摘要
The rapid proliferation of Generative AI (GenAI) has democratized the creation of hyper-realistic multimedia forgeries, posing severe threats to electronic Know Your Customer (eKYC) systems and digital forensic investigations. While visual synthesis has reached near-perfection, maintaining precise synchronization between lip movements (visemes) and speech signals (phonemes) remains a formidable challenge. To address this, we propose a novel Multimodal Deep Learning framework designed to detect high-fidelity Deepfakes by exploiting audio visual temporal inconsistencies. Beyond traditional feature fusion, our architecture integrates a Contrastive Synchronization Loss with a Transformer based Cross-Modal Attention mechanism. This hybrid objective explicitly enforces intra-class compactness for authentic pairs while amplifying the distance for asynchronous forgeries. Extensive experiments on FaceForensics++, DFDC, and a custom Vietnamese dataset (Vn-eKYC-Aug) demonstrate that our model achieves state-of-the-art performance, maintaining high robustness against video compression and environmental noise, though operational efficacy remains sensitive to extreme low-light conditions and diverse regional dialects. This research provides a resilient forensic layer for digital identity verification, ensuring evidence integrity in the GenAI era.
科研通智能强力驱动
Strongly Powered by AbleSci AI