Archive for SDEs

gradient flow for projected Langevin dynamics

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , on April 7, 2025 by xi'an

Daniel Lacker (Columbia U) gave a talk at the probability seminar of Paris Dauphine this week which I happened to attend by happenstance, on a recent paper, Projected Langevin dynamics and a gradient flow for entropic optimal transport, written with Giovanni Conforti and, Soumik Pal. The talk was quite progressive and I hence could follow most of it. The core idea is in studying Langevin-type diffusion dynamics that sample from an entropy-regularized optimal transport, i.e. looking for an optimal distribution (in the sense of achieving entropy minimisation problem within a Wasserstein space, with regularisation) obtained via a gradient flow equation (as eg in variational inference) that couples two SDEs that are recentred by conditional expectation terms. Expectations in the equations are estimated by a Nadaraya-Watson estimate in optimal transport problem (reminding me of SMC), with no theoretical derivation of an optimal bandwidth, and they achieve quantitive bounds on the convergence, namely for exponential convergence, energy decay and new logarithmic Sobolev inequalities. From the talk and a quick glance at the paper, it is unclear to me there are direct algorithmic consequences, since the SDEs need be discretised, while the expectation approximations are costly, being repeated at each iteration of the discretised SDE.

more than mostly MC

Posted in Kids, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , on December 27, 2024 by xi'an

The session of last Friday (and last one of 2024!) proved most interesting, with two (fully) Monte Carlo talks. The first one by Louis Grenioux was about an improvement on diffusion sampler, following a recent arXival by Maxime Noble and co-authors. Which brought me back to the long-standing multimodal challenge in Monte Carlo methods. When all modes of a target distribution are known, even roughly, this is not much of an issue since samplers can be arm-bent into visiting all these modes. But the problem becomes much harder when the location and a fortiori the number of modes are not known. The paper aims at adapting diffusion samplers towards a better exploration of the modes, albeit their location is known. Otherwise, using MCMC as a starting (reference) distribution would risk missing some of them. Which also explains why the authors can rely on the classical Gaussian mixtures proposal as a cheap substitute to neural networks (EBM). Since, within diffusion models, both intermediary distributions and their scores are intractable, they also introduce a variational parametric approximation that can be optimised.

The second talk was given by Guillaume Chennetier, in connection with his recent PhD thesis, developing a form of X-entropy sampling for rare events. As in nuclear plant major accidents. It took me a while to realise that PDMPs were not used as a simulation tool, as in the zigzag sampler and its avatars, The proposal involved creating a graph structure on the space of PDMP trajectories and designing the optimal importance process (yes, the one with zero variance!) using so-called committor functions that modify jump intensity and kernel, in a sequential way reminiscent of X-entropy. The approach recycles past trajectories as Monte Carlo elements if missing an adaptive mixture importance sampling (AMIS!) version that would bring more stability. The talk also included an interesting pointer to the availability of the distribution of the PDMP path, thus treated as a likelihood. (!). The signage at the entrance of the Monte-Carlo (mind the hyphen!) casino also made an appearance, reminding me of our memorable group picture on the same spot, eons ago! (But not of whom took the picture!)

mostly MCMC’s back

Posted in Statistics, University life with tags , , , , , , , , , , , , on September 13, 2024 by xi'an

Natural statistical science [#2]

Posted in Statistics with tags , , , , , , , , , , , , , , , on November 23, 2023 by xi'an

A rare occurrence of a Bayesian statistics paper in Nature with this “State estimation of a physical system with unknown governing equations” by Course and Nair. A variational Bayes modelling of a state system observed with noise, but without a physical model on the state (SDE) evolution itself. Which means a prior is set on a non-parametric or neural representation of the drift and a linear approximation is used for the variational approximation, leading to a Gaussian process as the approximate distribution. While this applies to highly complex models, like orbiting black holes, it is somewhat a surprise to meet this application of variational inference in a prestigious general science journal like Nature. (The picture above was taken on the train from Marseille at the end of the Bayes Fall school.)

“The approach is based on a technique called Bayesian inference, which is used widely, but which can be computationally challenging for complex systems.” B. Keith

séminaire parisien de statistique [09/01/23]

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , on January 22, 2023 by xi'an

I had missed the séminaire parisien de statistique for most of the Fall semester, hence was determined to attend the first session of the year 2023, the more because the talks were close to my interest. To wit, Chiara Amorino spoke about particle systems for McKean-Vlasov SDEs, when those are parameterised by several parameters, when observing repeatedly discretised versions, hereby establishing the consistence of a contrast estimator of these estimators. I was initially confused by the mention of interacting particles, since the work is not at all about related with simulation. Just wondering whether this contrast could prove useful for a likelihood-free approach in building a Gibbs distribution?

Valentin de Bortoli then spoke on diffusion Schrödinger bridges for generative models, which allowed me to better my understanding of this idea presented by Arnaud at the Flatiron workshop last November. The presentation here was quite different, using a forward versus backward explanation via a sequence of transforms that end up approximately Gaussian, once more reminiscent of sequential Monte Carlo. The transforms are themselves approximate Gaussian versions relying on adiscretised Ornstein-Ulhenbeck process, with a missing score term since said score involves a marginal density at each step of the sequence. It can be represented [as below] as an expectation conditional on the (observed) variate at time zero (with a connection with Hyvärinen’s NCE / score matching!) Practical implementation is done via neural networks.

Last but not least!, my friend Randal talked about his Kick-Kac formula, which connects with the one we considered in our 2004 paper with Jim Hobert. While I had heard earlier version, this talk was mostly on probability aspects and highly enjoyable as he included some short proofs. The formula is expressing the stationary probability measure π of the original Markov chain in terms of explorations between two visits to an accessible set C, more general than a small set. With at first an annoying remaining term due to the set not being Harris recurrent but which eventually cancels out. Memoryless transportation can be implemented because C is free for the picking, for instance the set where the target is bounded by a manageable density, allowing for an accept-reject step. The resulting chain is non-reversible. However, due to the difficulty to simulate from the target restricted to C, a second and parallel Markov chain is instead created. Performances, unsurprisingly, depend on the choice of C, but it can be adapted to the target on the go.