过度拟合
MNIST数据库
计算机科学
人工智能
分类器(UML)
卷积神经网络
特征向量
模式识别(心理学)
支持向量机
机器学习
反向传播
人工神经网络
作者
Sebastien Wong,Adam Gatt,V. Stamatescu,Mark D. McDonnell
标识
DOI:10.1109/dicta.2016.7797091
摘要
In this paper we investigate the benefit of augmenting data with synthetically created samples when training a machine learning classifier. Two approaches for creating additional training samples are data warping, which generates additional samples through transformations applied in the data-space, and synthetic over-sampling, which creates additional samples in feature-space. We experimentally evaluate the benefits of data augmentation for a convolutional backpropagation-trained neural network, a convolutional support vector machine and a convolutional extreme learning machine classifier, using the standard MNIST handwritten digit dataset. We found that while it is possible to perform generic augmentation in feature-space, if plausible transforms for the data are known then augmentation in data-space provides a greater benefit for improving performance and reducing overfitting.
科研通智能强力驱动
Strongly Powered by AbleSci AI