随机梯度下降算法
计算机科学
人工智能
进化算法
人口
进化计算
人工神经网络
深层神经网络
深度学习
随机优化
梯度下降
共同进化
互补性(分子生物学)
机器学习
数学优化
进化策略
健身景观
数学
生物
生态学
社会学
人口学
遗传学
作者
Xiaodong Cui,Zhang We,Zoltán Tüske,Michael Picheny
标识
DOI:10.48550/arxiv.1810.06773
摘要
We propose a population-based Evolutionary Stochastic Gradient Descent (ESGD) framework for optimizing deep neural networks. ESGD combines SGD and gradient-free evolutionary algorithms as complementary algorithms in one framework in which the optimization alternates between the SGD step and evolution step to improve the average fitness of the population. With a back-off strategy in the SGD step and an elitist strategy in the evolution step, it guarantees that the best fitness in the population will never degrade. In addition, individuals in the population optimized with various SGD-based optimizers using distinct hyper-parameters in the SGD step are considered as competing species in a coevolution setting such that the complementarity of the optimizers is also taken into account. The effectiveness of ESGD is demonstrated across multiple applications including speech recognition, image recognition and language modeling, using networks with a variety of deep architectures.
科研通智能强力驱动
Strongly Powered by AbleSci AI