判别式
卷积神经网络
计算机科学
人工智能
深度学习
稀缺
模式识别(心理学)
机器学习
数据建模
人工神经网络
班级(哲学)
深层神经网络
特征提取
语音识别
标记数据
环境数据
卷积(计算机科学)
训练集
任务分析
网络体系结构
声音(地理)
作者
Justin Salamon,Juan Pablo Bello
标识
DOI:10.1109/lsp.2017.2657381
摘要
The ability of deep convolutional neural networks (CNNs) to learn discriminative spectro-temporal patterns makes them well suited to environmental sound classification. However, the relative scarcity of labeled data has impeded the exploitation of this family of high-capacity models. This study has two primary contributions: first, we propose a deep CNN architecture for environmental sound classification. Second, we propose the use of audio data augmentation for overcoming the problem of data scarcity and explore the influence of different augmentations on the performance of the proposed CNN architecture. Combined with data augmentation, the proposed model produces state-of-the-art results for environmental sound classification. We show that the improved performance stems from the combination of a deep, high-capacity model and an augmented training set: this combination outperforms both the proposed CNN without augmentation and a “shallow” dictionary learning model with augmentation. Finally, we examine the influence of each augmentation on the model's classification accuracy for each class, and observe that the accuracy for each class is influenced differently by each augmentation, suggesting that the performance of the model could be improved further by applying class-conditional data augmentation.
科研通智能强力驱动
Strongly Powered by AbleSci AI