初始化
计算机科学
规范(哲学)
人工神经网络
人工智能
算法
应用数学
数学优化
数学
政治学
法学
程序设计语言
作者
Hongfei Yang,Xiaofeng Ding,Raymond Chan,Hui Hu,Yaxin Peng,Tieyong Zeng
摘要
Training deep neural networks can be difficult. For classical neural networks, the initialization method by Xavier and Yoshua which is later generalized by He, Zhang, Ren and Sun can facilitate stable training. However, with the recent development of new layer types, we find that the above mentioned initialization methods may fail to lead to successful training. Based on these two methods, we will propose a new initialization by studying the parameter space of a network. Our principal is to put constrains on the growth of parameters in different layers in a consistent way. In order to do so, we introduce a norm to the parameter space and use this norm to measure the growth of parameters. Our new method is suitable for a wide range of layer types, especially for layers with parameter-sharing weight matrices.
科研通智能强力驱动
Strongly Powered by AbleSci AI