Abstract
In this presentation, we will review how concepts from optimal transport can be applied to analyze various machine learning methods, particularly for sampling and training neural networks. The focus will be on using optimal transport to study dynamic flows in the space of probability distributions. The first example will focus on “flow-matching” sampling, which is based on the regression of advection fields. In its simplest case (diffusion models), this approach exhibits a gradient structure similar to that of optimal transport. We will then discuss Wasserstein gradient flows, where the flow minimizes a functional within the geometry of optimal transport. This framework allows us to model and understand the dynamics of probability distribution drift in neurons within two-layer networks. Finally, the last example will explore modeling the evolution of the probability distribution of tokens in deep transformer networks. This approach requires a modification of the optimal transport structure to incorporate the softmax normalization specific to attention mechanisms.