过度拟合
计算机科学
一般化
人工神经网络
频域
人工智能
过程(计算)
光学(聚焦)
深层神经网络
提前停车
机器学习
领域(数学分析)
算法
计算机视觉
操作系统
光学
物理
数学分析
数学
作者
Zhi‐Qin John Xu,Yaoyu Zhang,Yanyang Xiao
标识
DOI:10.1007/978-3-030-36708-4_22
摘要
Why deep neural networks (DNNs) capable of overfitting often generalize well in practice is a mystery [24]. To find a potential mechanism, we focus on the study of implicit biases underlying the training process of DNNs. In this work, for both real and synthetic datasets, we empirically find that a DNN with common settings first quickly captures the dominant low-frequency components, and then relatively slowly captures the high-frequency ones. We call this phenomenon Frequency Principle (F-Principle). The F-Principle can be observed over DNNs of various structures, activation functions, and training algorithms in our experiments. We also illustrate how the F-Principle helps understand the effect of early-stopping as well as the generalization of DNNs. This F-Principle potentially provides insight into a general principle underlying DNN optimization and generalization.
科研通智能强力驱动
Strongly Powered by AbleSci AI