
Archive for kernel Stein discrepancy descent
many folks for manifolds [computational methods for probability distributions on manifolds workshop]
Posted in pictures, Travel, University life with tags bootstrap, experimental design, France, generative model, Gromov-Wasserstein, group picture, IHP, Institut Henri Poincaré, inverse problems, kernel Stein discrepancy descent, manifold, manifold exploration, MCMC, mixture models, Paris, workshop on May 27, 2026 by xi'an
computational methods for probability distributions on manifolds (11-13 May, IHP, Paris)
Posted in Books, pictures, Statistics, Travel, University life with tags bootstrap, experimental design, generative model, Gromov-Wasserstein, IHP, Institut Henri Poincaré, inverse problems, kernel Stein discrepancy descent, manifold, manifold exploration, MCMC, mixture models, Paris, workshop on May 12, 2026 by xi'an
This week, we are running a small workshop on Computational methods for probability distributions on manifolds, whose size was dictated by the corresponding surface of the Institut Henri room allotted to us by the IHP administration. Very exciting theme and very exciting program, which more than make up for the unseasonal weather in Paris.
May 11
Guillaume Pouliot – MCMC on Manifolds in Economics
Alessandro Barp – Kernel and Stein discrepancies between distributions, à la Schwartz
Robin Ryder – Coupling MCMC on manifolds
Chang-Han Rhee – Experimental Design on Manifolds
May 12
Gilles Vilmart – High-order sampling of the invariant distribution of ergodic stochastic dynamics: preconditioning and postprocessing
Paul Breiding – Sampling from or near nonlinear algebraic varieties
Nick Whiteley – Statistical exploration of the Manifold Hypothesis
Judith Rousseau – Denoising diffusion Models under the Manifold Hypothesis : A dimension free convergence rate
Manon Michel – Convergence of non-reversible Markov processes via lifting and Flow Poincaré inequality
Tobias Grafke – Sampling Conditioned Diffusions via Pathspace Projected Monte Carlo
Miranda Holmes-Cerfon – Simulating sticky Brownian motion
Agnès Desolneux – Distances “à la Gromov-Wasserstein” for Gaussian Mixture Models
May 13
Giovanni Samaey – Multilevel interacting particle methods for sampling Bayesian inverse problems
Marylou Gabrié – Revisiting enhanced sampling driven by collective variables using generative models
Chris Walker – A Bayesian Perspective on the Maximum Score Problem
Lulu Kang – Active Learning for Manifold Gaussian Process Regression
scalable Monte Carlo for Bayesian learning [book review]
Posted in Books, Statistics, University life with tags Bayesian neural networks, book review, bouncy particle sampler, Cambridge University Press, CHANCE, Charles Stein, coordinate sampler, cup, HMC, IMS, IMS Monographs, kernel Stein discrepancy descent, Langevin diffusion, MALA, MCMC, monograph, Monte Carlo Statistical Methods, non-reversible MCMC, partly deterministic processes, PDMP, ULA on September 26, 2025 by xi'an
This book by Paul Fearnhead, Christopher Nemeth, Chris Oates, and Chris Sherlock is part of the IMS Monograph series. And published by Cambridge University Press. It covers most recent developments in MCMC methods, namely stochastic gradient MCMC (Chap. 3), non-reversible MCMC (Chap. 4), continuous-time MCMC (Chap. 5), and assessing and improving MCMC (Chap. 6). I find the book remarkable in its attention to rigour and clarity, without falling into overly technical derivations. It is perfectly suited for a graduate course to students with a solid mathematical background. In short, had I considered a new edition of our Monte Carlo Statistical Methods book to incorporate these advances, I could not done such a good job!
The first chapter provides a quick refresher of the background, from Monte Carlo principles, to Markov chains, SDEs, and the kernel “trick” (which requires a dozen pages of exposition). Nonetheless, it contains side remarks of true interest, including some suggestions I had not previously seen, as for instance an unusual introduction of the HMC algorithm as an underdamped Langevin diffusion. Chapter 2 prolongates this recap by covering reversible MCMC algorithms and the attached optimal scalings. This is done in a particularly friendly presentation that I intend to use in my own course. The HMC section is probably the best coverage I have seen on the topic, including most naturally the leapfrog steps.
Chapter 3 gets into stochastic gradient MCMC as an approximate MCMC, with nice arguments and formal convergence bounds. Again quite efficiently, if focussing almost solely on Gaussian settings (but including a neural network example). Similarly, Chapter 4 provides intuitive (if informal) arguments on the worth of non-reversible algorithms that are well-suited to a textbook of this level. This chapter introduces a PDMP sampler like the discrete bouncy particle sampler.
Chapter 5 is a (nicely) monstrous coverage of continuous time MCMC samplers that reaches very recent advances on PDMPs. The focus is on expressing them as limits, in order to derive mixing rates without extreme mathematical steps. (The chapter even includes a mention to the coordinate sampler that my PhD student Wu Changye derived in 2018!) Again a chapter I plan to use when teaching MCM methods, if possibly skipping some of the 66 pages.
Chapter 6 completes the monograph with a presentation of convergence assessment tools and diagnostics, exploiting the kernel trick, as well as convergence bounds that reflect very recent research in that domain. The conclusive section on optimal weights and optimal thinning will presumably be new to most readers. (Making me wonder if a link can be found with our importance Markov chain construct.)
[Disclaimer about potential self-plagiarism as usual: this post or an edited version will eventually appear in my Books Review section in CHANCE.]
Scalable Monte Carlo for Bayesian Learning [not yet a book review]
Posted in Books, Statistics, University life with tags 1⁰ North, Bayesian learning, book review, Cambridge University Press, continuous time MCMC, convergence diagnostics, cup, Gelman-Rubin statistic, Hamiltonian Monte Carlo, IMS Monographs, kernel Stein discrepancy descent, Markov chain Monte Carlo, MCMC, Metropolis adjusted Langevin algorithm, non-reversible MCMC, North, PDMP, piecewise deterministic, scalable Bayesian learning, scalable MCMC, stochastic differential equation, stochastic gradient MCMC on May 11, 2025 by xi'ansimulation as optimization [by kernel gradient descent]
Posted in Books, pictures, Statistics, University life with tags ABC, biking, Charles Stein, CREST, diffusions, discrepancies, Edo, Gare de Lyon, gradient descent, Hiroshige, INRIA, kernel Stein discrepancy descent, Kullback-Leibler divergence, maximum mean discrepancy, MCMC, Mokaplan, mollified discrepancy, New York city, One Hundred Famous Views of Edo, optimal transport, optimisation, Paris, simulation, SMC, Stein kernel on April 13, 2024 by xi'an
Yesterday, which proved an unseasonal bright, warm, day, I biked (with a new wheel!) to the east of Paris—in the Gare de Lyon district where I lived for three years in the 1980’s—to attend a Mokaplan seminar at INRIA Paris, where Anna Korba (CREST, to which I am also affiliated) talked about sampling through optimization of discrepancies.
This proved a most formative hour as I had not seen this perspective earlier (or possibly had forgotten about it). Except through some of the talks at the Flatiron Institute on Transport, Diffusions, and Sampling last year. Incl. Marilou Gabrié’s and Arnaud Doucet’s.
The concept behind remains attractive to me, at least conceptually, since it consists in approximating the target distribution, known up to a constant (a setting I have always felt standard simulation techniques was not exploiting to the maximum) or through a sample (a setting less convincing since the sample from the target is already there), via a sequence of (particle approximated) distributions when using the discrepancy between the current distribution and the target or gradient thereof to move the particles. (With no randomness in the Kernel Stein Discrepancy Descent algorithm.)
Ana Korba spoke about practically running the algorithm, as well as about convexity properties and some convergence results (with mixed performances for the Stein kernel, as opposed to SVGD). I remain definitely curious about the method like the (ergodic) distribution of the endpoints, the actual gain against an MCMC sample when accounting for computing time, the improvement above the empirical distribution when using a sample from π and its ecdf as the substitute for π, and the meaning of an error estimation in this context.
“exponential convergence (of the KL) for the SVGD gradient flow does not hold whenever π has exponential tails and the derivatives of ∇ log π and k grow at most at a polynomial rate”
