Archive for pseudo-marginal MCMC

BayesComp 2025.2

Posted in Kids, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on June 19, 2025 by xi'an


The main BayesComp²⁵ conference started with Pierre Jacob’s plenary talk on his recent advances on coupling for unbiased MCMC—currently ERC grantee on that topic—.  Raising lazy questions like using a different target or transition kernel for the second chain in the coupling, connecting the Poisson equation and control variates, handling the signed issue with the unbiased approximations. Interestingly, they obtain an unbiased estimator of the asymptotic variance of the unbiased estimator. And a correction for self-normalised importance sampling, which has some connections with our 1996 (?) pinball sampler. Also an evaluation of the median of means, rather than the average of means, which is a thing I had been (lazily) contemplating for a while  (On the greedy side, as I was writing my recovery exam for my Monte Carlo course, I realised the results Pierre presented could be somewhat recycled into exam problems!)

My first parallel session was on gradient-based methods with a talk by Francesca Crucinio on proximal particle Langevin algorithms (similar to the one she gave in PariSanté last year), a talk by Zhihao Wang on stereographic multiple try Metropolis(-Rosenbluth-Teller) that unsurprisingly recovers ergodicity thanks to the compactness of the ball. For which I wonder why a Normal proposal makes complete sense since one could consider a mover after the projection instead and why iid rather than repelling multiple proposals are used… The last speaker was just out from the plane from California, Siddharth Vishwanath who spoke about repelling-attracting HMC. With very nice animations of HMC, if reaching the main point of using both negative and positive frictions a few minutes before the session finished. The method preserves volume and potential, if not energy.

Speaking of which (energy), I find myself struggling with my less than 6 hours of sleep since arrival during the first afternoon session, despite a fiery hot spot lunch, which means in plainer terms that I alas dozed in and out of the talks. The second session saw Jack Jewson exposing in deeper details the exact PDMP algorithm for Gibbs measures  Jeremias Knoblauch mentioned yesterday. And Jonathan Huggins as well, using Gaussian processes as proxies for expected likelihoods, with lower guarantees than pseudo-marginal versions. In a mildly connected way, Robin Ryder went through the resolution of the ecological inference challenge they produce with Nicolas Chopin and Théo Valdoire (all authors with whom I am connected, Théo being a brillant student of our MASH Master last year and now in Harvard, hopefully till the end of his PhD!)

On the extra-academic curriculum, I had a yummy dinner in the Maxwell Hawker (street) food centre, incl. Xiao Long Bao that cooled down fast enough to avoid the usual scalding effect, plus rojak a mixed fruit and vegetable fried in a peanut sauce that I had never tasted before, popiah (ditto), chili noodles, and an appam with durian deepfried balls as a fabulous and unexpected dessert.

OWABI@BioInference2025 [29 May]

Posted in Mountains, Statistics, Travel, University life with tags , , , , , , , , , , , , , , on May 13, 2025 by xi'an

The next OWABI webinar is going to be quite special, consisting of two selected talks livestreamed from BioInference 2025, a conference on mathematical modelling and inference on (broadly speaking) biological system, taking place in Bardonecchia, Piedmont. The talks will take place on 29 May, 11am CEST (10am BST). The talks will be streamed on the OWABI MS Team Channel as usual.

1st OWABI Talk: 10-10.30am UK time

Speaker: Andrew Golightly (Durham University)

Title: Accelerating Bayesian inference for stochastic epidemic models using incidence data

Abstract: This work considers the case of performing Bayesian inference for stochastic epidemic compartment models, using incomplete time course data consisting of incidence counts that are either the number of new infections or removals in time intervals of fixed length. The most natural Markov jump process representation of the model is eschewed for reasons of computational efficiency, and replaced by a stochastic differential equation representation. This is further approximated to give a tractable Gaussian process, that is, the linear noise approximation (LNA). Unless the observation model linking the LNA to data is both linear and Gaussian, the observed data likelihood remains intractable. Unlike previous approaches that use the LNA in this setting, two approaches for marginalising over the latent process are considered: a correlated pseudo-marginal method and analytic marginalisation via a Gaussian approximation of the noise model. These approaches are compared using synthetic data with the best performing method applied to real data consisting of removal incidence of Oak Processionary moth nests in Richmond Park, London.

2nd OWABI Talk: 10.30-11am

Speaker: Henrik Häggström (Chalmers University)

Title: Simulation-based inference for stochastic nonlinear mixed-effects models with applications in systems biology

Abstract: We propose a novel methodology for Bayesian inference in hierarchical mixed-effects models. By building on our work, we construct a simulation-based inference (SBI) framework that is highly scalable, where amortized approximations to the likelihood and the parameters posterior are first obtained, and these are rapidly refined for each individual dataset, to ultimately approximate the parameters posterior across many individuals. Unlike the current state-of-art SBI methods, which use neural networks, our approximations are expressed via Gaussian mixture models, leading to easily trainable, parsimonious yet expressive surrogate models of both the likelihood function and the posterior distribution. The methodology is exemplified via stochastic differential equation mixed-effects models to describe translation kinetics after mRNA transfection, however the methodology is general and can accommodate other types of stochastic and deterministic models. We compare our approximate inference with exact pseudomarginal inference and show that our methodology is fast and competitive.

R[are]SS meeting

Posted in Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on September 29, 2024 by xi'an


Yesterday, I happened to be at the right time in the right place, as I was in Warwick for a RSS local section meeting on rare event simulation. (If missing the aurora borealis and the moon eclipse on previous nights!) And hence attended a seminar by Francesca Crucinio in six days!, as she talked about a turnkey approach to unbiased estimation of transforms of a moment, or wlog a mean μ, f(μ). A recent article with Nicolas Chopin (CREST) and Sumeet Singh, where they resort to Taylor expansions to achieve unbiasedness, using the Russian roulette trick to stop the summation from running to infinity. (As it happens, I heard Nicolas talk about this idea in the recent past namely at the ISBA-Fusion Sunday morn at Ca’Foscari.) Using a Taylor expansion is obviously natural and mathematically correct, albeit fraught with potential dangers [imho]:

  • the Taylor expansion involves central moments up to a random order R, which are harder & harder to estimate with increasing orders (i.e., more & more uncertain, with the possibility of infinite variance estimators after a certain order)
  • I did not spot a discussion on the moment estimators, that seems to rely on k iid replicas for the k-th moment
  • a lot of calibration ensues, from the choice of the centre x⁰ to the (artificial) distribution of the stopping value R, to the parameterisation of the random variable attached to the moment μ
  • the paper insists on recycling simulations to stabilise the moment estimators and ensure consistency, as a primary level of Rao-Blackwellisation, but this only applies to the smallest order moments and could be devised in many different ways, with varying computing costs
  • consistency of the estimate is not necessarily needed, as for instance for pseudo-marginal applications
  • as often with Russian roulette, positive quantities may receive negative estimations that are dominated by truncations to the positive real line (and alternating series offer the use of sandwiching estimators)
  • for the above reason, it is not always reasonable to tunnel vision on unbiasedness and alternative estimates like bridge sampling solutions could be integrating towards improving the quality of the estimator (especially since the conditions for finite variance involve unknown quantities)
  • while f-Taylored solutions like harmonic mean estimators for f(x)=1/x are not necessarily a panacea, they could be included in the comparison or as control variates

The first talk by Mathias Rousset was investigating adaptive multilevel sampling, a form of nested sampler, at the theoretical level, while the third talk by Tobias Grafke was a repetition of a talk he gave at the masterclass the interface between computational physics and computational statistics, last April.

Saddlepoint Monte Carlo and its application to vote transfers

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , on January 3, 2024 by xi'an


Our former Dauphine Master student Théo Voldoire (now PhD-ing at Harvard), along with Nicolas Chopin, Guillaume Rateau, and Robin Ryder (now at Imperial), arXived a month ago a paper with title Saddlepoint Monte Carlo and its Application to Exact Ecological Inference, essentially the outcome of his Master thesis last summer. Nicolas came to present the paper at our Mostly MC seminar. The motivating example is about vote transfers (to surviving candidates for a second round, as in the French presidential and deputorial elections) based on results from polling stations across rounds, which equates filling a contingency table with known margins, exploiting multiple tools like characteristic functions, inverse Fourier transform, its Monte Carlo version, pseudo-marginal MCMC, tilting, exponential families, quasi Monte Carlo! Which also reminded me of the time Reuven Rubinstein was occasionally visiting Paris, defending the cross entropy approach. Among multiple questions raised by this original approach to an “old” problem, one may think of the model misspecification issue that political analysts would not fail to raise, namely that the transfer estimates are based on multinomial models, that all models are wrong, &tc. We discussed briefly about this during the seminar, the suggestion being a predictive check by cross validation. The talk also brought to mind highly probable applications to privacy, and possibly to capture recapture.

Exact MCMC with differentially private moves

Posted in Statistics with tags , , , , , , , on September 25, 2023 by xi'an

“The algorithm can be made differentially private while remaining exact in the sense that its target distribution is the true posterior distribution conditioned on the private data (…) The main contribution of this paper arises from the simple  observation that the penalty algorithm has a built-in noise in its calculations which is not desirable in any other context but can be exploited for data privacy.”

Another privacy paper by Yldirim and Ermis (in Statistics and Computing, 2019) on how MCMC can ensure privacy. For free. The original penalty algorithm of Ceperley and Dewing (1999) is a form of Metropolis-Hastings algorithm where the Metropolis-Hastings acceptance probability is replaced with an unbiased estimate (e.g., there exists an unbiased and Normal estimate of the log-acceptance ratio, λ(θ, θ’), whose exponential can be corrected to remain unbiased).  In that case, the algorithm remains exact.

“Adding noise to λ(θ, θ) may help with preserving some sort of data privacy in a Bayesian framework where [the posterior], hence λ(θ, θ), depends on the data.”

Rather than being forced into replacing the Metropolis-Hastings acceptance probability with an unbiased estimate as in pseudo-marginal MCMC, the trick here is in replacing λ(θ, θ’) with a Normal perturbation, hence preserving both the target (as shown by Ceperley and Dewing (1999)) and the data privacy, by returning a noisy likelihood ratio. Then, assuming that the difference sensitivity function for the log-likelihood [the maximum difference c(θ, θ’) over pairs of observations of the difference between log-likelihoods at two arbitrary parameter values θ and θ’] is decreasing as a power of the sample size n, the penalty algorithm is differentially private, provided the variance is large enough (in connection with c(θ, θ’)] after a certain number of MCMC iterations. Yldirim and Ermis (2019) show that the setting covers the case of distributed, private, data. even though the efficiency decreases with the number of (protected) data silos. (Another drawback is that the data owners must keep exchanging likelihood ratio estimates.