人工智能
计算机科学
不变(物理)
培训(气象学)
事件(粒子物理)
算法
语音识别
数学
模式识别(心理学)
排列(音乐)
训练集
声音(地理)
Boosting(机器学习)
集合(抽象数据类型)
噪音(视频)
作者
Baofu Duan,Yinhuan Dong,Tughrul Arslan,John Thompson
标识
DOI:10.1109/icassp55912.2026.11462845
摘要
Sound event localization and detection (SELD) requires simultaneous recognition of events and estimation of their spatial positions. Existing output representations have notable limitations. Single-branch heads couple detection and localization too tightly, so improving one task inevitably affects others. Conversely, conventional multi-branch heads fail to represent overlapping events of the same class, reducing their effectiveness. To overcome these limitations, we propose TriAD, a tri-head output representation for SELD. It combines track-wise localization with Auxiliary Duplicating Permutation Invariant Training (ADPIT), enabling the capture of overlapping events. Experiments on DCASE2025 stereo dataset show that TriAD achieves state-of-the-art F1 score performance. We further conduct the first systematic study of gradient-aware multi-task optimization for SELD and demonstrate that, among all tested strategies, Projected Conflicting Gradient (PCGrad) delivers the largest F1 score gains. These findings highlight the importance of robust representations and effective optimization strategies for advancing SELD systems.
科研通智能强力驱动
Strongly Powered by AbleSci AI