卷积神经网络
计算机科学
循环神经网络
声音(地理)
人工神经网络
人工智能
语音识别
声学
物理
作者
Bastian Estay Zamorano,Ali Dehghan Firoozabadi,Alessio Brutti,Pablo Adasme,David Zabala‐Blanco,Pablo Palacios Játiva,César A. Azurdia-Meza
出处
期刊:Electronics
[Multidisciplinary Digital Publishing Institute]
日期:2025-07-10
卷期号:14 (14): 2778-2778
标识
DOI:10.3390/electronics14142778
摘要
Sound event localization and detection (SELD) is a fundamental task in spatial audio processing that involves identifying both the type and location of sound events in acoustic scenes. Current SELD models often struggle with low signal-to-noise ratios (SNRs) and high reverberation. This article addresses SELD by reformulating direction of arrival (DOA) estimation as a multi-class classification task, leveraging deep convolutional recurrent neural networks (CRNNs). We propose and evaluate two modified architectures: M-DOAnet, an optimized version of DOAnet for localization and tracking, and M-SELDnet, a modified version of SELDnet, which has been designed for joint SELD. Both modified models were rigorously evaluated on the STARSS23 dataset, which comprises 13-class, real-world indoor scenes totaling over 7 h of audio, using spectrograms and acoustic intensity maps from first-order Ambisonics (FOA) signals. M-DOAnet achieved exceptional localization (6.00° DOA error, 72.8% F1-score) and perfect tracking (100% MOTA with zero identity switches). It also demonstrated high computational efficiency, training in 4.5 h (164 s/epoch). In contrast, M-SELDnet delivered strong overall SELD performance (0.32 rad DOA error, 0.75 F1-score, 0.38 error rate, 0.20 SELD score), but with significantly higher resource demands, training in 45 h (1620 s/epoch). Our findings underscore a clear trade-off between model specialization and multifunctionality, providing practical insights for designing SELD systems in real-time and computationally constrained environments.
科研通智能强力驱动
Strongly Powered by AbleSci AI