| 标题 |
Mutual-guidance framework for audio DeepFake detection via multi-dimensional feature interaction |
| 网址 | |
| DOI | |
| 其它 | With the rapid advancement of text-to-speech (TTS) and voice cloning (VC) technologies, the perceptual quality of synthetic speech has approached that of natural speech, posing substantial challenges to conventional auditory-based detection methods. To address this issue, a novel Mutual-Guidance Framework (MGF) is introduced, designed to integrate both physical and high-level semantic representations for enhanced detection accuracy. In this framework, Mel-Frequency Cepstral Coefficients (MFCCs) and Linear-Frequency Cepstral Coefficients (LFCCs) are employed to capture low-level acoustic characteristics, while a pre-trained Wav2Vec encoder is utilized to extract semantic embeddings, thereby constructing a multi-level representational hierarchy. To achieve precise cross-modal alignment and deep interaction, a Hierarchical Cross-Attention Fusion (HCAF) module is incorporated, enabling multi-level information exchange between physical and semantic features. |
| 求助人 | |
| 下载 |
PDF的下载单位、IP信息已删除
(2025-6-4)