声音(地理)
声源定位
计算机科学
声学
事件(粒子物理)
估计
语音识别
物理
工程类
量子力学
系统工程
作者
Nao Sato,Masahiro Yasuda,Shoichiro Saito,Noboru Harada
标识
DOI:10.1109/icassp49660.2025.10888141
摘要
Sound Event Localization and Detection (SELD) is the combined task of detecting sound events and estimating their spatial locations. We propose a Sound source Distance Estimation (SDE) method for SELD that utilizes a physics-informed prior. The conventional data-driven approach of SDE for SELD can handle complex situations where sound sources move or overlap, thanks to multitask learning. Deep Neural Network (DNN)-based SELD systems are generally pre-trained with synthesized data to compensate for the lack of real data. However, in the context of recent SELD tasks, the performance of DNN models, which are pre-trained with synthetic data, significantly degrades on real data. One possible cause of this is that the DNN models overfit sound characteristics that do not exist in the real data, i.e., are not physically reasonable. Therefore, we propose a hybrid SDE method that utilizes a physics-informed prior in a data-driven SELD system. Focusing on the specificity of the expected sound PoWer Level (PWL) of the sound sources depending on the class, we set the typical PWL for each class as a prior. To improve real-world applicability, we also adopt the sound attenuation model as a physics-informed prior for explicitly utilizing physical laws in SDE. Experimental results suggested the effectiveness of utilizing a physics-informed prior in SDE for SELD to improve its applicability to the real world.
科研通智能强力驱动
Strongly Powered by AbleSci AI