过参数化设定下动量随机梯度下降的指数收敛速率

Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting

Mathematical Programming · 2026
被引 0
ABS 4

中文导读

证明了动量随机梯度下降在非凸优化中,对于任意固定超参数,在满足局部Polyak-Łojasiewicz不等式和过参数化方差假设下,具有指数收敛速率,并分析了摩擦参数的最优选择。

Abstract

Abstract We prove explicit bounds on the exponential rate of convergence for the momentum stochastic gradient descent scheme (MSGD) for arbitrary, fixed hyperparameters (learning rate, friction parameter) and its continuous-in-time counterpart in the context of non-convex optimization. The results are shown for objective functions satisfying a local Polyak-Łojasiewicz inequality and under assumptions on the variance of MSGD that are satisfied in overparametrized settings. Moreover, we analyze the optimal choice of the friction parameter and show that the MSGD process almost surely converges to a local minimum.

非凸优化随机梯度下降动量方法过参数化