Archive for HMC
January session of the mostly Monte Carlo seminar (16/01, 3pm)
Posted in Statistics, University life with tags #ERCSyG, HMC, Monte Carlo Statistical Methods, mostly Monte Carlo seminar, multi-armed bandits, Ocean, Paris, PariSanté campus, Porte de Versailles, privacy, seminar, stereographic MCMC on January 9, 2026 by xi'anscalable Monte Carlo for Bayesian learning [book review]
Posted in Books, Statistics, University life with tags Bayesian neural networks, book review, bouncy particle sampler, Cambridge University Press, CHANCE, Charles Stein, coordinate sampler, cup, HMC, IMS, IMS Monographs, kernel Stein discrepancy descent, Langevin diffusion, MALA, MCMC, monograph, Monte Carlo Statistical Methods, non-reversible MCMC, partly deterministic processes, PDMP, ULA on September 26, 2025 by xi'an
This book by Paul Fearnhead, Christopher Nemeth, Chris Oates, and Chris Sherlock is part of the IMS Monograph series. And published by Cambridge University Press. It covers most recent developments in MCMC methods, namely stochastic gradient MCMC (Chap. 3), non-reversible MCMC (Chap. 4), continuous-time MCMC (Chap. 5), and assessing and improving MCMC (Chap. 6). I find the book remarkable in its attention to rigour and clarity, without falling into overly technical derivations. It is perfectly suited for a graduate course to students with a solid mathematical background. In short, had I considered a new edition of our Monte Carlo Statistical Methods book to incorporate these advances, I could not done such a good job!
The first chapter provides a quick refresher of the background, from Monte Carlo principles, to Markov chains, SDEs, and the kernel “trick” (which requires a dozen pages of exposition). Nonetheless, it contains side remarks of true interest, including some suggestions I had not previously seen, as for instance an unusual introduction of the HMC algorithm as an underdamped Langevin diffusion. Chapter 2 prolongates this recap by covering reversible MCMC algorithms and the attached optimal scalings. This is done in a particularly friendly presentation that I intend to use in my own course. The HMC section is probably the best coverage I have seen on the topic, including most naturally the leapfrog steps.
Chapter 3 gets into stochastic gradient MCMC as an approximate MCMC, with nice arguments and formal convergence bounds. Again quite efficiently, if focussing almost solely on Gaussian settings (but including a neural network example). Similarly, Chapter 4 provides intuitive (if informal) arguments on the worth of non-reversible algorithms that are well-suited to a textbook of this level. This chapter introduces a PDMP sampler like the discrete bouncy particle sampler.
Chapter 5 is a (nicely) monstrous coverage of continuous time MCMC samplers that reaches very recent advances on PDMPs. The focus is on expressing them as limits, in order to derive mixing rates without extreme mathematical steps. (The chapter even includes a mention to the coordinate sampler that my PhD student Wu Changye derived in 2018!) Again a chapter I plan to use when teaching MCM methods, if possibly skipping some of the 66 pages.
Chapter 6 completes the monograph with a presentation of convergence assessment tools and diagnostics, exploiting the kernel trick, as well as convergence bounds that reflect very recent research in that domain. The conclusive section on optimal weights and optimal thinning will presumably be new to most readers. (Making me wonder if a link can be found with our importance Markov chain construct.)
[Disclaimer about potential self-plagiarism as usual: this post or an edited version will eventually appear in my Books Review section in CHANCE.]
BayesComp 2025.2
Posted in Kids, pictures, Statistics, Travel, University life with tags appam, Arianna Rosenbluth, BayesComp 2025, Bayesian neural networks, Bayesian robustness, Bayesian semi-parametrics, characteristic function, chili pepper, Chinatown, cut models, durian, equator, exam, Gaussian processes, Gibbs measure, Gibbs posterior, Hamiltonian Monte Carlo, Harvard University, HMC, MASH, Maxwell Hawker market, National University Singapore, NUS, PDMP, Poisson equation, popiah, pseudo-marginal MCMC, rojak, safe Bayes, self-normalised importance sampling, Shanghai, Singapore, stereographic MCMC, stochastic gradient MCMC, street food, unbiased MCMC, Université Paris Dauphine, xiaolongbao, zigzag algorithm on June 19, 2025 by xi'an
The main BayesComp²⁵ conference started with Pierre Jacob’s plenary talk on his recent advances on coupling for unbiased MCMC—currently ERC grantee on that topic—. Raising lazy questions like using a different target or transition kernel for the second chain in the coupling, connecting the Poisson equation and control variates, handling the signed issue with the unbiased approximations. Interestingly, they obtain an unbiased estimator of the asymptotic variance of the unbiased estimator. And a correction for self-normalised importance sampling, which has some connections with our 1996 (?) pinball sampler. Also an evaluation of the median of means, rather than the average of means, which is a thing I had been (lazily) contemplating for a while (On the greedy side, as I was writing my recovery exam for my Monte Carlo course, I realised the results Pierre presented could be somewhat recycled into exam problems!)
My first parallel session was on gradient-based methods with a talk by Francesca Crucinio on proximal particle Langevin algorithms (similar to the one she gave in PariSanté last year), a talk by Zhihao Wang on stereographic multiple try Metropolis(-Rosenbluth-Teller) that unsurprisingly recovers ergodicity thanks to the compactness of the ball. For which I wonder why a Normal proposal makes complete sense since one could consider a mover after the projection instead and why iid rather than repelling multiple proposals are used… The last speaker was just out from the plane from California, Siddharth Vishwanath who spoke about repelling-attracting HMC. With very nice animations of HMC, if reaching the main point of using both negative and positive frictions a few minutes before the session finished. The method preserves volume and potential, if not energy.
Speaking of which (energy), I find myself struggling with my less than 6 hours of sleep since arrival during the first afternoon session, despite a fiery hot spot lunch, which means in plainer terms that I alas dozed in and out of the talks. The second session saw Jack Jewson exposing in deeper details the exact PDMP algorithm for Gibbs measures Jeremias Knoblauch mentioned yesterday. And Jonathan Huggins as well, using Gaussian processes as proxies for expected likelihoods, with lower guarantees than pseudo-marginal versions. In a mildly connected way, Robin Ryder went through the resolution of the ecological inference challenge they produce with Nicolas Chopin and Théo Valdoire (all authors with whom I am connected, Théo being a brillant student of our MASH Master last year and now in Harvard, hopefully till the end of his PhD!)

On the extra-academic curriculum, I had a yummy dinner in the Maxwell Hawker (street) food centre, incl. Xiao Long Bao that cooled down fast enough to avoid the usual scalding effect, plus rojak a mixed fruit and vegetable fried in a peanut sauce that I had never tasted before, popiah (ditto), chili noodles, and an appam with durian deepfried balls as a fabulous and unexpected dessert.
JSM 2024, Portland, Day 2
Posted in pictures, Running, Statistics, Travel, University life with tags American Statistical Association, ASA, conference centre, confidence distribution, COPSS Elizabeth L. Scott Award, data science, doubly intractable posterior, evolutionary Monte Carlo, fusion, gene expression, hidden networks, HMC, Joint Statistical Meeting, JSM 2024, Mount Hamilton, Mount Hood National Forest, open water swimming, Oregon, PDMP, Portland, ranger station, SNIP, unification, United States of America, US Department of Agriculture, Willamette River, Zigzag, zigzag algorithm on August 7, 2024 by xi'an
By happenstance, I started my day in the cybersecurity session. With (again) hardly a soul in the room… A first talk on avoiding herding and achieving asymptotic truth learning (about a binary outcome) in a graph by putting constraints on the graph structure, without any clear connection with statistics or cybersecurity. Even less for the second talk on optimising masks. Only with the third one came cybersecurity motivations, the focus being on a two-player Stackelberg game already used in this framework. The result proper was about estimating the parameter of a (rather unrealistic) posited model reproducing the adversarial actions. The last talk about jailbreak attacks was again off-field by miles.

Then attended (as intended!) the 2024 Blackwell-Rosenbluth Award session, featuring the nominees Sharmistha Guha (Texas A&M) on multiple network inference, Simon Mak (Duke) on using Bayesian surrogate models, Guanyang Wang (Rutgers) who recently spoke at our mostly Monte Carlo seminar, Akihiko Nishimura (John Hopkins) on a unification of HMC and PDMPs, most appropriate when located next to Mount Hamilton and the Zigzag river!, and Maria Skoularidou (MIT) on evolutionary Monte Carlo for gene expression. Alas with hardly anyone in the room.

In the afternoon, I went to the (very well-attended this time!, with no seat available for many attendees, incl. yours truly!) COPSS Elizabeth L. Scott Lecture by my friend from Rutgers, Regina Liu, on the highly relevant challenge of combining inferences from diverse data sources. Using the (definitely Rutgerian!) approach of confidence distributions!
Last (late) afternoon, I went swimming from Kevin Duckworth dock, just below the conference centre. Water was quite warm (and green), with a few other swimmers, and no stomachical after-effect, so far. Hence I returned there once again this afternoon.
repelling-attracting Hamiltonian Monte Carlo
Posted in Books, pictures, Statistics, University life with tags computing time, friction, Hamiltonian Monte Carlo, HMC, multimodality, repelling-attracting, Stanford University on June 25, 2024 by xi'an
Lasrt week, Siddharth Vishwanath and Hyungsuk Tak—whom I first met at an MCQMC session about multimodal sampling at MCqMC 2016 in Stanford, same year as my San Fran’ half-marathon race, most memorable of all my races!)—proposed a Repelling-Attracting Hamiltonian Monte Carlo (raHMC) algorithm, towards sampling from multimodal distributions.
“The success of raHMC for sampling from multimodal distributions crucially hinges on the choice of the [three] tuning parameter[s]”
The concept behind raHMC is to endow an HMC algorithm with an added friction term that slows down moves, except it can get turned into an acceleration effect when the friction coefficient γ becomes negative. In a proposal remindful of leapfrog half-time moves, raHMC proceeds by switching the sign of this coefficient γ half-way of an artificial time parameter T that is representing the inter-simulation time between two successive states of the Markov chain. By this aggregation of opposite forces, the resulting algorithm satisfies the detailed-balance condition, hence is reversible and preserves symplectic structure and volume. If not energy.
“…a direct application of the repelling-attracting mechanism to NUTS may not be straightforward. Lastly, we have not been able to guarantee that raHMC conserves energy”
Given the dependence on the tuning parameters, I fear implementing the algorithm may prove delicate in more complex settings, e.g. when the number of modes is unknown, as the acceleration component is rather blind to the actual target. In addition, the leapfrog integrator may prove quite slow in low density regions, which are visited about half the time.
“This, however, comes at the price of a higher computational cost, as the auto-tuning procedure for raHMC tends to favor longer trajectories, and therefore requires more gradient evaluations per step”
Numerical experiments show, indeed, that the algorithm is much slower than others, as this occurence of a 8.5s execution time for HMC vs a corresponding 1094s for raHM…
