Séminaire Images Optimisation et Probabilités
(Maths-IA) What is the long-run behavior of stochastic gradient descent? A large deviations analysis
Franck Iutzeler
( Institut de Mathématiques de Toulouse )Salle de Conférences
le 09 janvier 2025 à 11:15
Abstact: We examine the long-run distribution of stochastic gradient descent (SGD) in general, non-convex problems. Specifically, we seek to understand which regions of the problem's state space are more likely to be visited by SGD, and by how much. Using an approach based on the theory of large deviations and randomly perturbed dynamical systems, we show that the long-run distribution of SGD resembles the Boltzmann-Gibbs distribution of equilibrium thermodynamics with temperature equal to the method's step-size and energy levels determined by the problem's objective and the statistics of the noise. Joint work w/ W. Azizian, J. Malick, P. Mertikopoulos
https://arxiv.org/abs/2406.09241 published at ICML 2024