09:30 to 11:00
Lecture

Score derived from denoising and generalization

Stéphane Mallat
Amphithéâtre Marguerite de Navarre, Site Marcelin Berthelot
Open to all, subject to availability
-

Abstract

Thanks to Tweedie’s identity, a neural network can estimate the score of a probability distribution by computing a denoising estimator that minimizes the mean squared error. This is done by training a neural network that takes as input a noisy image with varying levels of noise and computes a denoising process that produces the minimum error. Deep neural networks appear capable of computing an estimator close to the optimal estimator on complex data, and can therefore compute the score. This score is used in the inverse diffusion equation, which generates new data from a white noise realization.

Given the quality of the images generated by score diffusion, one might wonder whether this is simply a matter of memorizing the training data used to train the denoising network. Numerical experiments show that this is the case as long as the training dataset is not sufficiently large. However, these experiments also demonstrate that when the dataset is sufficiently large, the network generalizes, and the score diffusion algorithm generates new data that is independent of the training data.

A fundamental challenge is to understand how a neural network can compute an optimal denoising estimator despite the curse of high dimensionality. The course first considers optimal linear denoising using the Wiener estimator. A connection is established between the score of Gaussian distributions and the Wiener estimator. In the case of a stationary process, Wiener filtering is a convolution filter with a specified Fourier transform.

Events