人工智能
计算机科学
光学(聚焦)
模式识别(心理学)
特征(语言学)
事件(粒子物理)
卷积神经网络
过程(计算)
监督学习
特征提取
班级(哲学)
特征学习
融合
传感器融合
目标检测
训练集
深度学习
标记数据
相似性(几何)
高分辨率
计算机视觉
特征向量
显著性图
作者
Xichang Cai,Liangxiao Zuo,Ziyi Liu,Menglong Wu,Hongyang Guo,Xuejing Sun
标识
DOI:10.1109/caibda65784.2025.11182885
摘要
The data required for training a supervised Sound Event Detection (SED) model needs to be manually annotated, which makes acquiring it quite difficult. During the process of obtaining temporal localization from weakly supervised networks using Class Activation Map (CAM), due to the low resolution of CAM, the accuracy of the positioning is poor. In this paper, we propose a weakly supervised SED method based on a deep-shallow feature fusion CNN. It extracts feature maps from multiple convolutional layers to capture the network's focus areas on features, and constructs a contrastive loss for training to make the network focus on features more accurately. Furthermore, aggregating the CAMs extracted from different convolutional layers with weighted aggregation further improves the performance of the method proposed in this paper. Experiments on the Domestic environment sound event detection (DESED) dataset demonstrate that the proposed method achieves promising performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI