事件(粒子物理)
计算机科学
人工智能
计算机视觉
噪音(视频)
估计
模式识别(心理学)
声音(地理)
信号处理
算法
声源定位
信号(编程语言)
作者
Qingjing Wan,Ying Hu,Hao Huang
标识
DOI:10.1109/icassp55912.2026.11462435
摘要
Conventional sound event localization and detection (SELD) focuses on sound event detection (SED) and direction of arrival (DOA) estimation, which provide limited spatial information. The 3D SELD task addresses this limitation by incorporating source distance estimation (SDE). In this paper, we propose a similarity-guided aggregation network (SGA-Net) for sound event localization and detection with source distance estimation (3D SELD) that combines the spectrum-based branch and pre-trained model-based branch. We design a multi-scale feature extraction (MSFE) module for the spectrum-based branch to capture rich acoustic information and a similarity-guided feature fusion (SGFF) module that includes a spatial feature transform (SFT) module for effective fusion of intermediate features from both branches. The experimental results demonstrate that our proposed SGA-Net achieves the best performance in all metrics on the STARSS23 dataset, verifying the effectiveness of SGA-Net. The main code will be available at https://github.com/QingJWan/SGANet.
科研通智能强力驱动
Strongly Powered by AbleSci AI