Archive for Markov kernel

mostly Monte Carlo in June

Posted in Statistics, University life with tags , , , , , , , , , , , , , , , , , , on May 30, 2026 by xi'an

The last episode of the academic year for our mostly Monte Carlo seminar, next week:

On Friday 05/06/26, from 3-5pm at PariSanté Campus

15h00: Sam Livingstoke (University College London)

Skew-symmetric numerical schemes for stochastic differential equations: strong convergence and multi-level extension
I will discuss recent work fusing together two strands of the applied mathematics and statistics literature, one concerned with developing flexible probability distributions for data that rely on a small number of parameters, and another concerned with developing numerical integration schemes to simulate stochastic processes.  The specific case that I will focus on uses the skew-symmetric family of probability distributions introduced by Adelchi Azzalini and co-authors to approximate the transition kernels of diffusion processes over small time steps, producing alternative numerical schemes to the classical Euler-Maruyama approach.  Applying the scheme to the overdamped Langevin diffusion leads to an unadjusted version of the Barker proposal Metropolis-Hastings algorithm.  In earlier work weak accuracy was established over finite and infinite time scales, crucially without needing a globally Lipschitz assumption on the drift of the stochastic differential equation.  I will review this and then discuss more recent work establishing strong convergence in the mean-squared sense using a novel coupling between the numerical and exact processes.  This also enables the development of a multi-level Monte Carlo scheme, which I will discuss the merits of with particular focus on the superlinear drift case, as compared to Euler and Tamed Euler alternatives.
This is joint work with Yuga Iguchi, Giorgos Vasdekis & Rui-Yang Zhang.
16h00: Dana Naderi (Université Paris Dauphine PSL)
Approximating evidence via bounded harmonic means

Efficient Bayesian model selection relies on the model evidence or marginal likelihood, whose computation often requires evaluating an intractable integral. The harmonic mean estimator (HME) has long been a standard method of approximating the evidence. While computationally simple, the version introduced by Newton and Raftery (1994) potentially suffers from infinite variance. To overcome this issue, Gelfand and Dey (1994) defined a standardized representation of the estimator based on an instrumental function and Robert and Wraith (2009) later proposed to use higher posterior density (HPD) indicators as instrumental functions. Following this approach, a practical method is proposed, based on an elliptical covering of the HPD region with non-overlapping ellipsoids. The resulting estimator, called the Elliptical Covering Marginal Likelihood Estimator (ECMLE), not only eliminates the infinite-variance issue of the original HME and allows exact volume computations, but is also able to be used in multimodal settings. Through several examples, we illustrate that ECMLE outperforms other recent methods such as THAMES and its improved version (Metodiev et al. 2025). Moreover, ECMLE demonstrates lower variance a key challenge that subsequent HME variants have sought to address-and provides more stable evidence approximations, even in challenging settings.

This is joint work with Kaniav Kamari, Dareen Wraith & myself (X).

mostly Monte Carlo [13/03]

Posted in Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on March 10, 2026 by xi'an

A new episode of our mostly Monte Carlo seminar, very soon coming near you (if in Paris):

On Friday 13/02/26, from 3-5pm at PariSanté Campus

15h00: Pierre Del Moral (INRIA, Bordeaux)

On the Kantorovich contraction of Markov semigroup

We present a novel operator theoretic framework to study the contraction properties of Markov semigroups with respect to a general class of Kantorovich semi-distances, which notably includes Wasserstein distances. This rather simple contraction cost framework combines standard Lyapunov techniques with local contraction conditions. Our results can be applied to both discrete time and continuous time Markov semigroups, and we illustrate their wide applicability in the context of (i) Markov transitions on models with boundary states, including bounded domains with entrance boundaries, (ii) operator products of a Markov kernel and its adjoint, including two-block-type Gibbs samplers, (iii) iterated random functions and (iv) diffusion models, including overdampted Langevin diffusion with convex at infinity potentials.

16h00: Bob Carpenter (Flatiron Institute, New York)

GIST, WALNUTS, and Continuous Nutpie: mass-matrix and step-size adaptation for Hamiltonian Monte Carlo

I will introduce Gibbs self tuning (GIST), our new technique for coupling tuning parameters and conditionally Gibbs-sampling them per iteration in Hamiltonian Monte Carlo. Then I will turn to the within-orbit adaptive NUTS (WALNUTS) sampler, which adapts the step size every leapfrog step in order to conserve the Hamiltonian. Empirical evaluations on varying multi-scale target distributions, including Neal’s funnel and the Stock-Watson stochastic volatility time-series model, demonstrate that WALNUTS achieves substantial improvements in sampling efficiency and robustness. I will review the Nutpie mass-matrix adaptation scheme, which is designed to minimize Fisher divergence by estimating the mass matrix as the geometric midpoint (aka barycenter) between the inverse covariance of the draws and the covariance of the scores of the draws. Then I will describe a continuously adapting version that adapts per iteration by continuously discounting the past rather than updating in fixed blocks. I will also show how the Adam optimizer outperforms dual averaging for step-size adaptation. I will conclude by considering a lock-free multi-threading implementation that automatically monitors adaptation and sampling for convergence for automatic stopping.

important and published [Markov chains]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , on February 26, 2024 by xi'an

Our importance Markov chain paper (with Charly Andral (PhD, Paris Dauphine), Randal Douc, and Hugo Marival (PhD, Telecom SudParis) has been published (on-line) by stochastic processes and their applications (SPA). Incidentally, my first publication in this journal. To paraphrase its abstract, it sort of bridges the (unsuspected?) gap between rejection sampling and importance sampling, moving from one to the other through a tuning parameter. Based on a modified sample of an instrumental Markov chain targeting an instrumental distribution (typically via a MCMC kernel), rather than the target of interest, the Importance Markov chain produces an extended Markov chain whose (first) marginal distribution converges to said target distribution. For instance, when targeting a multimodal distribution, the instrumental distribution can be chosen as a tempered version of the target and this frees the algorithm to explore its multiple modal regions more efficiently. We also derive a Law of Large Numbers and a Central Limit Theorem as well as prove geometric ergodicity for this extended kernel under mild assumptions on the instrumental kernel. Computationally, the algorithm is easy to implement and preexisting librairies can be used to sample from the instrumental distribution.

important Markov chains

Posted in Books, Statistics, University life with tags , , , , , , , , , , , on July 21, 2022 by xi'an

With Charly Andral (PhD, Paris Dauphine), Randal Douc, and Hugo Marival (PhD, Telecom SudParis), we just arXived a paper on importance Markov chains that merges importance sampling and MCMC. An idea already mentioned in Hastings (1970) and even earlier in Fodsick (1963), and later exploited in Liu et al.  (2003) for instance. And somewhat dual of the vanilla Rao-Backwellisation paper Randal and I wrote a (long!) while ago. Given a target π with a dominating measure π⁰≥Mπ, using a Markov kernel to simulate from this dominating measure and subsampling by the importance weight ρ does produce a new Markov chain with the desired target measure as invariant distribution. However, the domination assumption is rather unrealistic and a generic approach can be implemented without it, by defining an extended Markov chain, with the addition of the number N of replicas as the supplementary term… And a transition kernel R(n|x) on N with expectation ρ, which is a minimal(ist) assumption for the validation of the algorithm.. While this initially defines a semi-Markov chain, an extended Markov representation is also feasible, by decreasing N one by one until reaching zero, and this is most helpful in deriving convergence properties for the resulting chain, including a CLT.  While the choice of the kernel R is free, the optimal choice is associated with residual sampling, where only the fractional part of ρ is estimated by a Bernoulli simulation.

null recurrent = zero utility?

Posted in Books, R, Statistics with tags , , , , , , , on April 28, 2022 by xi'an

The stability result that the ratio

\dfrac{\sum^T_{t=1} f(\theta^{(t)})}{\sum^T_{t=1} g(\theta^{(t)})}\qquad(1)

converges holds for a Harris π-null-recurrent Markov chain for all functions f,g in L¹(π) [Meyn & Tweedie, 1993, Theorem 17.3.2] is rather fascinating. However, it is unclear it can be useful in simulation environments, as for the integral priors we have been studying over the years with Juan Antonio Cano and Diego Salmeron Martinez. Above, the result of an experiment where I simulated a Markov chain as a Normal random walk in dimension one, hence a Harris π-null-recurrent Markov chain for the Lebesgue measure λ, and monitored the stabilisation of the ratio (1) when using two densities for f and g,  to its expected value (1, shown by a red horizontal line). There is quite a variability in the outcome (repeated 100 times),  but the most intriguing is the quick stabilisation of most cumulated averages to values different from 1. Even longer runs display this feature

which I would blame on the excursions of the random walk far away from the central regions for both f and g, that is on long sequences where zeroes keep being added to numerator and denominators in (1). As far as integral approximation is concerned, this is not very helpful!