Source Localization Using Distributed Microphones in Reverberant Environments Based on Deep Learning and Ray Space Transform

混响虚假关系计算机科学稳健性（进化）参数统计话筒声源定位非线性系统代表（政治）多向性算法空格（标点符号）人工智能声学模式识别（心理学）数学物理声波机器学习节点（物理）电信量子力学操作系统统计生物化学化学法学基因声压政治政治学

作者

Luca Comanducci,Federico Borra,Paolo Bestagini,Fabio Antonacci,Stefano Tubaro,Augusto Sarti

出处

期刊：IEEE/ACM transactions on audio, speech, and language processing [Institute of Electrical and Electronics Engineers]
日期：2020-01-01 卷期号：28: 2238-2251 被引量：21

标识

DOI：10.1109/taslp.2020.3011256

摘要

In this article we present a methodology for source localization in reverberant environments from Generalized Cross Correlations (GCCs) computed between spatially distributed individual microphones. Reverberation tends to negatively affect localization based on Time Differences of Arrival (TDOAs), which become inaccurate due to the presence of spurious peaks in the GCC. We therefore adopt a data-driven approach based on a convolutional neural network, which, using the GCCs as input, estimates the source location in two steps. It first computes the Ray Space Transform (RST) from multiple arrays. The RST is a convenient representation of the acoustic rays impinging on the array in a parametric space, called Ray Space. Rays produced by a source are visualized in the RST as patterns, whose position is uniquely related to the source location. The second step consists of estimating the source location through a nonlinear fitting, which estimates the coordinates that best approximate the RST pattern obtained through the first step. It is worth noting that training can be accomplished on simulated data only, thus relaxing the need of actually deploying microphone arrays in the acoustic scene. The localization accuracy of the proposed techniques is similar to the one of SRP-PHAT, however our method demonstrates an increased robustness regarding different distributed microphones configurations. Moreover, the use of the RST as an intermediate representation makes it possible for the network to generalize to data unseen during training.

求助该文献

最长约 10秒，即可获得该文献文件

Source Localization Using Distributed Microphones in Reverberant Environments Based on Deep Learning and Ray Space Transform

今日热心研友