Archive for HMC

mostly M[ay]C

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , , , , on May 22, 2024 by xi'an

With the details from the second speaker:

Adaptive MCMC sampling using a Metropolized PDMP sampler combined with a No-U-Turn criterion

Augustin Chevallier, Université de Strasbourg

Adaptivity in MCMC algorithms is hard to achieve. In Hamiltonian Monte Carlo, for example, it is possible to tune the path length using the No-U-Turn sampler, but the numerical step size cannot be adapted; it can only be tuned. We propose here a new class of algorithm based on Metropolizing a numerical approximation of a PDMP sampler. Like HMC, these samplers require two parameters: a numerical step size and a path length. Unlike HMC, both parameters can be adapted. This paves the way for more robust sampling algorithms, especially for difficult target densities.

 

 

repulsive sampling

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , on January 31, 2024 by xi'an

After a long absence from the monthly Séminaire Parisien de Statistique I attended one today at IHP, including a talk by Diala Hawat on repelled point processes for numerical integration by Hawat et al. The goal is to get (and prove) a universal variance improvement for numerical integration by applying a form of determinantal processes to initial simulations, as eg iid (Poisson process) sampling (without accounting for the O(N²) cost in moving these points). The repelled points are obtain by a single (why single?) move based on a force function (as shown in the slide below), inspired by a Coulomb potential (in the sense that said move appears as one gradient step along the potential). Which reminded me of the pinball sampler, even though the inverse norm was just there to create infinite repulsion near each point. A surprising feature of this repelling step is that it even modifies a (QMC) Sobol process with also an (empirical) improvement in the variance. I wonder if one could construct an MCMC algorithm that would target a joint distribution, maybe via a copula representation, maybe via an equivalent version of HMC.


As an aside, the Bakhvalov results on the existence of a worst case integrand for any deterministic or random sequence (see top slide) made me wonder what the shape of this worst case function is, esp. for a QMC sequence (eg, Sobol). And whether or not they are of any relevance as a counterfactor to the optimal importance functions.

re-MCM’d

Posted in Statistics, University life with tags , , , , , , , , , , , , , , , on June 28, 2023 by xi'an


When I entered the classroom on the Jussieu campus where the Monday morning session on PDMP was taking place, some friends told me a badge was waiting for me at the registration desk of MCM 2023! Most surprisingly since I had received a deregistration message a few weeks earlier. Great PDMP session with non-reversible tempering (warning: some PDMPs were harmed in the tempering process), scaling PDMPs for highly anisotropic targets, a PDMP form of reversible jump (with sticky floors!) and a more adaptive version of no-U turn via… PDMPs. After lunch at the nearby Grande Mosquée de Paris (I had not visited since… 1974!), I attended the slice sampler session, where mileage varied imho. With a take on doubly-intractable targets that did not seem to relate to the existing literature.

Approximation Methods in Bayesian Analysis [#2]

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on June 22, 2023 by xi'an

A more theoretical Day #2 of the workshop, with Debdeep Pati comparing two representations of Gaussian processes with significantly different efficiencies, and Aad van der Vaart presenting a form of linearisation for a range of inverse problems, Kolyan Ray debiasing Lasso impacts by variational Bayes, although through a somewhat intricate process that distanced the procedure from Bayesian grounds imho, Judith Rousseau (Dauphine) also drifting away from Bayesian canons by looking anew at empirical Bayes with surprising differences from genuine B analysis, connecting with the cutoff phenomenon she and Kerrie exhibited in their 2011 mixture paper, as well as labelling the marginal likelihood a misspecified model. Trevor Campbell and Sinead Williamson both provided Bayesian perspectives on normalising flows, in particular the impact of computer imprecision on reversibility, leading to the notion of shadow paths (screenshot below), while Giovanni Rebaudo talked about mixtures supported by trees, a fascinating object!

On Day #3, Marc Beaumont talked on a mixture of composite likelihood à la Ryden, making me wonder of optimisation of blocks for HMC? EP-ABC, with the issue of the unknown amount of approximation, and adaptivity?, Maria de Iorio presented work on finite and infinite mixtures with repulsive (Coulomb) priors, achieving a unified framework, plus known evidence (?), with a correlated talk by Federico Camerlenghi in the afternoon, with novel notions (for me) of Palm measures and calculus, and another correlated talk by María-Fernanda Gil Leyva Villa, on stick-breaking processes for species sampling with dependent length variables, with related improvements in Gibbs implementation (screenshot below).
This was followed by two theoretical talks on continuous time processes by Paul Jenkins (Warwick) on the fine properties of the Flemming-Viot process, with mentions of Don Dawson’s results reminding me of the 1988 and 1989 summers I spent at Carleton University, where he was located at the time, and Matteo Ruggieri, with the novel (to me) notion of dual Markov processes that could prove useful in a lot of latent variable models. Fabrizio Leisen expanded on his early work on partial exchangeability and Steve MacEachern on dependent quantile pyramids, which relate to quantile regression, a constant source of puzzlement for me. Motivating the perspective by robustness and misspecification arguments. But I am a wee bit puzzled by the distinction between quantile pyramids and other non-parametric solutions.

On the outdoor front (in early mornings), choppy waters at sea (in the Sugiton calanque, pictured above) thanks to the endless mistral wind, nice run down from Mont Puget with friends, limited utility of my rented mountain bike (except to reach the nearest supermarket, 3km away)

sampling, transport, and diffusions

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , on November 18, 2022 by xi'an


This week, I am attending a very cool workshop at the Flatiron Institute (not in the Flatiron building!, but close enough) on Sampling, Transport, and Diffusions, organised by Bob Carpenter and Michael Albergo. It is quite exciting as I do not know most participants or their work! The Flatiron Institute is a private institute focussed on fundamental science funded by the Simons Foundation (in such working conditions universities cannot compete with!).

Eric Vanden-Eijden gave an introductory lecture on using optimal transport notion to improve sampling, with a PDE/ODE approach of continuously turning a base distribution into a target (formalised by the distribution at time one). This amounts to solving a velocity solution to an KL optimisation objective whose target value is zero. Velocity parameterised as a deep neural network density estimator. Using a score function in a reverse SDE inspired by Hyvärinnen (2005), with a surprising occurrence of Stein’s unbiased estimator, there for the same reasons of getting rid of an unknown element. In a lot of environments, simulating from the target is the goal and this can be achieved by MCMC sampling by normalising flows, learning the transform / pushforward map.

At the break, Yuling Yao made a very smart remark that testing between two models could also be seen as an optimal transport, trying to figure an optimal transform from one model to the next, rather than the bland mixture model we used in our mixtestin paper. At this point I have no idea about the practical difficulty of using / inferring the parameters of this continuum but one could start from normalising flows. Because of time continuity, one would need some driving principle.

Esteban Tabak gave another interest talk on simulating from a conditional distribution, which sounds like a no-problem when the conditional density is known but a challenge when only pairs are observed. The problem is seen as a transport problem to a barycentre obtained as a distribution independent from the conditioning z and then inverting. Constructing maps through flows. Very cool, even possibly providing an answer for causality questions.

Many of the transport talks involved normalizing flows. One by [Simons Fellow] Christopher Jazynski about adding to the Hamiltonian (in HMC) an artificial flow field  (Vaikuntanathan and Jarzynski, 2009) to make up for the Hamiltonian moving too fast for the simulation to keep track. Connected with Eric Vanden-Eijden’s talk in the end.

An interesting extension of delayed rejection for HMC by Chirag Modi, with a manageable correction à la Antonietta Mira. Johnatan Niles-Weed provided a nonparametric perspective on optimal transport following Hütter+Rigollet, 21 AoS. With forays into the Sinkhorn algorithm, mentioning Aude Genevay’s (Dauphine graduate) regularisation.

Michael Lindsey gave a great presentation on the estimation of the trace of a matrix by the Hutchinson estimator for sdp matrices using only matrix multiplication. Solution surprisingly relying on Gibbs sampling called thermal sampling.

And while it did not involve optimal transport, I gave a short (lightning) talk on our recent adaptive restore paper: although in retrospect a presentation of Wasserstein ABC could have been more suited to the audience.