Archive for PariSanté campus

mostly Monte Carlo in June

Posted in Statistics, University life with tags , , , , , , , , , , , , , , , , , , on May 30, 2026 by xi'an

The last episode of the academic year for our mostly Monte Carlo seminar, next week:

On Friday 05/06/26, from 3-5pm at PariSanté Campus

15h00: Sam Livingstoke (University College London)

Skew-symmetric numerical schemes for stochastic differential equations: strong convergence and multi-level extension
I will discuss recent work fusing together two strands of the applied mathematics and statistics literature, one concerned with developing flexible probability distributions for data that rely on a small number of parameters, and another concerned with developing numerical integration schemes to simulate stochastic processes.  The specific case that I will focus on uses the skew-symmetric family of probability distributions introduced by Adelchi Azzalini and co-authors to approximate the transition kernels of diffusion processes over small time steps, producing alternative numerical schemes to the classical Euler-Maruyama approach.  Applying the scheme to the overdamped Langevin diffusion leads to an unadjusted version of the Barker proposal Metropolis-Hastings algorithm.  In earlier work weak accuracy was established over finite and infinite time scales, crucially without needing a globally Lipschitz assumption on the drift of the stochastic differential equation.  I will review this and then discuss more recent work establishing strong convergence in the mean-squared sense using a novel coupling between the numerical and exact processes.  This also enables the development of a multi-level Monte Carlo scheme, which I will discuss the merits of with particular focus on the superlinear drift case, as compared to Euler and Tamed Euler alternatives.
This is joint work with Yuga Iguchi, Giorgos Vasdekis & Rui-Yang Zhang.
16h00: Dana Naderi (Université Paris Dauphine PSL)
Approximating evidence via bounded harmonic means

Efficient Bayesian model selection relies on the model evidence or marginal likelihood, whose computation often requires evaluating an intractable integral. The harmonic mean estimator (HME) has long been a standard method of approximating the evidence. While computationally simple, the version introduced by Newton and Raftery (1994) potentially suffers from infinite variance. To overcome this issue, Gelfand and Dey (1994) defined a standardized representation of the estimator based on an instrumental function and Robert and Wraith (2009) later proposed to use higher posterior density (HPD) indicators as instrumental functions. Following this approach, a practical method is proposed, based on an elliptical covering of the HPD region with non-overlapping ellipsoids. The resulting estimator, called the Elliptical Covering Marginal Likelihood Estimator (ECMLE), not only eliminates the infinite-variance issue of the original HME and allows exact volume computations, but is also able to be used in multimodal settings. Through several examples, we illustrate that ECMLE outperforms other recent methods such as THAMES and its improved version (Metodiev et al. 2025). Moreover, ECMLE demonstrates lower variance a key challenge that subsequent HME variants have sought to address-and provides more stable evidence approximations, even in challenging settings.

This is joint work with Kaniav Kamari, Dareen Wraith & myself (X).

incoming mostly Monte Carlo [14 April, PariSanté campus]

Posted in pictures, Statistics, University life with tags , , , , , , , , , , , , , , , on April 9, 2026 by xi'an

The next Mostly Monte Carlo seminar will be this very Friday, 10/04/26, at PariSanté Campus. With Shiva Darshan and Pierre Monmarché speaking on the following topics:
15h: Shiva Darshan Maximal-reflection couplings on manifolds: some specific examples
Explicit Markovian couplings can be used to build Markov Chain Monte Carlo methods such unbiased MCMC or coupling based control variates. For sampling from probability measures supported on Euclidean space, one typically uses a synchronous coupling, a maximal-reflection coupling (also known as a discrete-time sticky coupling), or some variant of the two. For probability measures supported on Riemannian manifolds, the situation is less clear cut. While the Kendall-Cranston coupling of Brownian motions on manifolds has been successfully applied in theoretical works, it is ill-suited for building explicit algorithms. In this talk, we will discuss some of the obstacles to extending Euclidean maximal-reflection couplings to manifolds and present some special cases for which these obstacles can be easily overcome. With applications to Stereographic MCMC in mind, we detail particular couplings of random walks on the sphere.
16h: Pierre Monmarché A post-sampling reweighting method for multi-modal target measures
Even when the modes are identified and sampled locally with MCMC methods, a difficulty to sample multi-modal measures is to correctly estimate the relative probabilities of each of these modes, which requires to observe many transitions between them (which are rare events). We will present an approach based on variational inference which exploits the local samples, aiming only at estimating the relative weights between them. When the modes are well separated, this amount to some entropy estimations.

Bob’s talk at PariSanté

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , on March 25, 2026 by xi'an

We had a wonderful time (and an unusually large audience) at the mostly Monte Carlo seminar last week as Pierre del Moral and Bob Carpenter both presented on exciting recent developments of theirs! Pierre talked about Kantorovich contraction of Markov semigroups, which sounds rather daunting!, but actually covers fairly general and generic convergence results, using tools like potentials and Lyapunov contractions, reminding me of the early days of MCMC and the papers of Gareth Roberts (University of Warwick), Jeff Rosenthal, Richard Tweedie and others.

Bob then spoke about the latest version of NUTS, the within-orbit adaptive NUTS (WALNUTS) sampler, which adapts the step size at every leapfrog step in order to conserve the Hamiltonian and keep the path stable enough. The adaptation is facilitated by incorporating this step size as an extra parameter with an attached distribution, that the authors call Gibbs self tuning (GIST), for coupling tuning parameters and conditionally Gibbs-sampling them per iteration in Hamiltonian Monte Carlo. This has been done in the past, incl. in some of my papers (e.g., Andrieu & Robert, 2004), but I could not cite a particular reference during the seminar.

Further light reflections that came to mind during Bob’s talk:

  • with NUTS, if cycling is feasible in a finite time, we could wait for a second passage at the starting point and then get back halfway (with the difficulty of detecting this second passage)
  • changing the kinetic matrix at each leapfrog jump is actually Riemannian HMC (and with cubic cost!)
  • the doubling mechanism in both the original NUTS and in biased progressive NUTS is simulation wasting
  • but so is (surprise, surprise!) finding adaptive mass matrices for WALNUTS at reasonable costs

mostly Monte Carlo [13/03]

Posted in Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on March 10, 2026 by xi'an

A new episode of our mostly Monte Carlo seminar, very soon coming near you (if in Paris):

On Friday 13/02/26, from 3-5pm at PariSanté Campus

15h00: Pierre Del Moral (INRIA, Bordeaux)

On the Kantorovich contraction of Markov semigroup

We present a novel operator theoretic framework to study the contraction properties of Markov semigroups with respect to a general class of Kantorovich semi-distances, which notably includes Wasserstein distances. This rather simple contraction cost framework combines standard Lyapunov techniques with local contraction conditions. Our results can be applied to both discrete time and continuous time Markov semigroups, and we illustrate their wide applicability in the context of (i) Markov transitions on models with boundary states, including bounded domains with entrance boundaries, (ii) operator products of a Markov kernel and its adjoint, including two-block-type Gibbs samplers, (iii) iterated random functions and (iv) diffusion models, including overdampted Langevin diffusion with convex at infinity potentials.

16h00: Bob Carpenter (Flatiron Institute, New York)

GIST, WALNUTS, and Continuous Nutpie: mass-matrix and step-size adaptation for Hamiltonian Monte Carlo

I will introduce Gibbs self tuning (GIST), our new technique for coupling tuning parameters and conditionally Gibbs-sampling them per iteration in Hamiltonian Monte Carlo. Then I will turn to the within-orbit adaptive NUTS (WALNUTS) sampler, which adapts the step size every leapfrog step in order to conserve the Hamiltonian. Empirical evaluations on varying multi-scale target distributions, including Neal’s funnel and the Stock-Watson stochastic volatility time-series model, demonstrate that WALNUTS achieves substantial improvements in sampling efficiency and robustness. I will review the Nutpie mass-matrix adaptation scheme, which is designed to minimize Fisher divergence by estimating the mass matrix as the geometric midpoint (aka barycenter) between the inverse covariance of the draws and the covariance of the scores of the draws. Then I will describe a continuously adapting version that adapts per iteration by continuously discounting the past rather than updating in fixed blocks. I will also show how the Adam optimizer outperforms dual averaging for step-size adaptation. I will conclude by considering a lock-free multi-threading implementation that automatically monitors adaptation and sampling for convergence for automatic stopping.

Bayesian, adversarial, oceanic, privacy

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , on March 6, 2026 by xi'an

We just arXived a new paper on Bayesian privacy! We meaning Cameron Bell, Antoine Luciano, Timothy Johnston and myself, as members of my ERC OCEAN lab at PariSanté and Paris Dauphine. While sharing the same ground as my recent paper with James Bailie, Joshua Bon and Judith Rousseau, this one is definitely more mainstream Bayesian in that the entire decision process falls under the Bayesian hat, with the ultimate decision being the choice of the release mechanism by the data holder (or hoarder!). To rationalise this decision process, we break the framework as resulting from the actions of three actors, namely the data holder, Alice, the data scientist, Bob, and the eavesdropper. Eve. (As in my earlier posts on solving Le Monde’s math puzzles, we could have used pronouns from other cultures, but I feared this would have confused some of the readers. Incidentally, I found out that the earliest use of the first two pronouns was within the groundbreaking cryptography 1977 paper of Rivest, Shamir and Adleman, bringing the RSA algorithm to the World! With Eve appearing in an early, highly-cited privacy paper by Montréal’s Bennett, Brassard, and (unconnected to me!) Robert, in 1988.)

We thus consider a Bayesian setting in which, given data x, held by Alice, inference is to be performed by Bob on a parameter θ. Performing such inference requires Alice releasing information derived from x, which may contain sensitive content, exploited by Eve. Our approach is to compare Alice’s release mechanisms according to both the quality of inference on θ (from Bob’s viewpoint) and the privacy leakage regarding x (sought by Eve and dreaded by Alice). To formalise this evaluation, we posit that Alice refers to a loss function that is a linear combination of Bob’s and Eve’s losses, the weight on Eve’s loss being then negative. (An alternative to be considered in future work is Alice using a ratio of Bob’s and Eve’s losses, possibly set to different powers, the rationale being that a zero loss for Eve is intolerable for Alice.) As in Bayesian experimental design, a prior on the data is necessary for Eve to infer on the hidden data based on the release mechanism and released output and for Alice to evaluate the risk of said release mechanism . (They may differ, as long as they are both made public.) To calibrate Alice’s loss, we opted for a balance that returns the same risk for a full data release and a total lack of release. In specific, informed, settings, other weights could be chosen. While finding the optimal release strategy is impossible but for highly discrete settings, the framework obviously allows for the ranking of natural strategies like insufficient statistics and synthetic datasets. Comments welcome!