动量(技术分析)
李普希茨连续性
趋同(经济学)
二次方程
随机梯度下降算法
数学优化
应用数学
计算机科学
数学
正多边形
数学分析
人工智能
经济
人工神经网络
几何学
财务
经济增长
作者
Yanli Liu,Yuan Gao,Wotao Yin
标识
DOI:10.48550/arxiv.2007.07989
摘要
SGD with momentum (SGDM) has been widely applied in many machine learning tasks, and it is often applied with dynamic stepsizes and momentum weights tuned in a stagewise manner. Despite of its empirical advantage over SGD, the role of momentum is still unclear in general since previous analyses on SGDM either provide worse convergence bounds than those of SGD, or assume Lipschitz or quadratic objectives, which fail to hold in practice. Furthermore, the role of dynamic parameters has not been addressed. In this work, we show that SGDM converges as fast as SGD for smooth objectives under both strongly convex and nonconvex settings. We also establish \textit{the first} convergence guarantee for the multistage setting, and show that the multistage strategy is beneficial for SGDM compared to using fixed parameters. Finally, we verify these theoretical claims by numerical experiments.
科研通智能强力驱动
Strongly Powered by AbleSci AI