计算机科学
混响
声源定位
路径(计算)
话筒
麦克风阵列
人工智能
噪音(视频)
相(物质)
维数(图论)
人工神经网络
模式识别(心理学)
声学
数学
声音(地理)
物理
声压
电信
图像(数学)
量子力学
程序设计语言
纯数学
作者
Bing Yang,Hong Liu,Xiaofei Li
标识
DOI:10.48550/arxiv.2202.07859
摘要
Multiple moving sound source localization in real-world scenarios remains a challenging issue due to interaction between sources, time-varying trajectories, distorted spatial cues, etc. In this work, we propose to use deep learning techniques to learn competing and time-varying direct-path phase differences for localizing multiple moving sound sources. A causal convolutional recurrent neural network is designed to extract the direct-path phase difference sequence from signals of each microphone pair. To avoid the assignment ambiguity and the problem of uncertain output-dimension encountered when simultaneously predicting multiple targets, the learning target is designed in a weighted sum format, which encodes source activity in the weight and direct-path phase differences in the summed value. The learned direct-path phase differences for all microphone pairs can be directly used to construct the spatial spectrum according to the formulation of steered response power (SRP). This deep neural network (DNN) based SRP method is referred to as SRP-DNN. The locations of sources are estimated by iteratively detecting and removing the dominant source from the spatial spectrum, in which way the interaction between sources is reduced. Experimental results on both simulated and real-world data show the superiority of the proposed method in the presence of noise and reverberation.
科研通智能强力驱动
Strongly Powered by AbleSci AI