Abstract
Among recent advances in machine learning , one of the most impressive is undoubtedly generative AI, which makes it possible to create increasingly realistic samples of sounds, images, and videos from a finite set of examples. At the heart of this revolution are diffusion models, which use the log-likelihood gradient within the framework of stochastic differential equations to generate new samples.
In this talk, we will briefly introduce diffusion models before analyzing the dynamics of generation in a well-controlled high-dimensional case: the mixture of two Gaussians. Using methods from statistical physics, we will analytically demonstrate that the generation of new data by a score model based on the empirical distribution involves various transitions. First, we will identify a “speciation” transition, during which the fate of the sample is sealed and its class can no longer be changed. This speciation is then followed by a collapse (or memorization) transition, after which the trajectory is irrevocably drawn toward one of the data points in the training set in order to reproduce it exactly.
The theoretical conclusions we will establish for a Gaussian mixture model will then be generalized to arbitrary distributions and validated on realistic datasets.