Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting
证明了动量随机梯度下降在非凸优化中,对于任意固定超参数,在满足局部Polyak-Łojasiewicz不等式和过参数化方差假设下,具有指数收敛速率,并分析了摩擦参数的最优选择。
Abstract We prove explicit bounds on the exponential rate of convergence for the momentum stochastic gradient descent scheme (MSGD) for arbitrary, fixed hyperparameters (learning rate, friction parameter) and its continuous-in-time counterpart in the context of non-convex optimization. The results are shown for objective functions satisfying a local Polyak-Łojasiewicz inequality and under assumptions on the variance of MSGD that are satisfied in overparametrized settings. Moreover, we analyze the optimal choice of the friction parameter and show that the MSGD process almost surely converges to a local minimum.