MaLGa logoMaLGa black extendedMaLGa white extendedUniGe ¦ MaLGaUniGe ¦ MaLGaUniversita di Genova | MaLGaUniversita di GenovaUniGe ¦ EcoSystemics
Seminar

The Long-Run Behavior of SGD in Non-Convex Landscapes: A Large Deviation Analysis

13/05/2026

azizian - [Naso Bocca Sorriso]

Title

The Long-Run Behavior of SGD in Non-Convex Landscapes: A Large Deviation Analysis


Speaker

Waïss Azizian - Université Grenoble Alpes


Abstract

Stochastic gradient descent (SGD) has been around for more than 70 years, yet it remains the workhorse for training today's largest machine learning systems, from LLMs to reinforcement learning and recommender models. Despite this long history and overwhelming success, one fundamental question about its behaviour on general objectives remains open: on non-convex objectives, what is the long-run distribution of SGD? In particular, which regions of the parameter space are visited the most by SGD – and by how much?

Answering this is difficult because the landscape of the loss function can be extraordinarily complex, especially in modern deep networks. In this work, our goal is to give a general, principled answer for arbitrary non-convex objectives.

Using large deviations and the theory of randomly perturbed dynamical systems, we show that the long-run distribution of SGD takes the form of a Boltzmann–Gibbs law: the step-size acts as a temperature, and the “energy levels'' are determined jointly by the objective and the noise. As a consequence, SGD concentrates exponentially around the minimum-energy state — which need not be the global minimum — while other regions are visited with probabilities exponentially proportional to their energy.

Finally, we will briefly touch on a complementary question — how long will it take SGD to reach a desirable region, such as the global minimum? Using similar tools, we obtain matching upper and lower bounds showing that this time is controlled by the most “costly’’ obstacles in the landscape, linking global geometry with the statistics of the noise.

Taken together, these works provide a complete characterization of the long-run behavior of SGD on general non-convex objectives, opening the door to a principled understanding of the optimization dynamics of modern machine learning systems.


Bio

Waïss Azizian is a final-year PhD student at Université Grenoble Alpes, France, advised by Franck Iutzeler, Jérôme Malick, and Panayotis Mertikopoulos. His research focuses on stochastic optimization, with a focus on applications in machine learning and deep learning, drawing on tools from dynamical systems, probability, and large deviations theory. Before his PhD, he studied at ENS Paris and graduated from the MVA master.


When

Wednesday, May 13th, 11:30


Where

Room 216 - DIMA/DIBRIS, via Dodecaneso 35, Genova