光谱图
计算机科学
语音识别
合成数据
语音合成
语音活动检测
信号(编程语言)
时域
语音处理
变压器
人工智能
模式识别(心理学)
工程类
计算机视觉
电压
电气工程
程序设计语言
作者
Amit Kumar Singh Yadav,Kratika Bhagtani,Sriram Baireddy,Paolo Bestagini,Stefano Tubaro,Edward J. Delp
出处
期刊:
日期:2024-03-18
卷期号:: 11171-11175
被引量:1
标识
DOI:10.1109/icassp48485.2024.10446471
摘要
With recent advancements in generating synthetic speech, tools to generate high-quality synthetic speech impersonating any human speaker are easily available. Several incidents report misuse of high-quality synthetic speech for spreading misinformation and for large-scale financial frauds. Many methods have been proposed for detecting synthetic speech; however, there is limited work on localizing the synthetic segments within the speech signal. In this work, our goal is to localize the synthetic speech segments in a partially synthetic speech signal. Most existing methods for synthetic speech localization obtain features from either the time domain waveform or the spectrogram representation of the speech signal. In this work, we propose Multi-Domain ResNet Transformer (MDRT) that obtains multi-domain features from both the time domain and the spectrogram representation of a speech signal to localize synthetic speech segments. MDRT uses transformer neural networks to obtain multi-domain features and processes them using a ResNet-style neural network. We use the PartialSpoof dataset to examine the performance of MDRT on localizing synthetic speech segments of varying duration. Our results show that MDRT performs better than several existing synthetic speech localization methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI