Archive for importance sampling

mostly Monte Carlo [09/10, PSC]

Posted in Statistics, University life with tags , , , , , , , , , , , , on October 4, 2026 by xi'an

The next episode of our mostly Monte Carlo seminar is next Friday (9 October) at PariSanté Campus (room #8) with speakers

15:00 – Víctor Elvira, University of Edinburgh

16:00 – Edoardo Bandoni, Université Paris Dauphine-PSL

Víctor Elvira, “Rethinking self-normalized importance sampling”

Self-normalized importance sampling (SNIS) is one of the most widely used Monte Carlo techniques for inference with unnormalized target distributions. Despite its usefulness, SNIS is often viewed simply as a normalized version of ordinary importance sampling, and many of its methodological questions remain largely unexplored. In this talk, we revisit SNIS from a unified perspective. We first introduce a generalized formulation of self-normalized importance sampling based on coupled proposals, showing that the classical SNIS estimator is only one member of a broader family of Monte Carlo estimators with new opportunities for variance reduction. We then consider the classical SNIS estimator and present adaptive algorithms that learn proposals tailored to its optimal proposal distribution, together with theoretical guarantees including consistency, asymptotic normality, and convergence of the proposal. Together, these developments suggest that self-normalized importance sampling should be regarded as a distinct Monte Carlo methodology, with its own theory, optimality principles, and algorithmic design.

Edoardo Bandoni, “Rate-Optimal Randomised Kernel Quadrature”

Kernel quadrature is widely used to approximate integrals of smooth functions, with the worst-case error typically decaying at the minimax rate n-α/d for smoothness α in dimension d. Existing rate-optimal methods often depend on deterministic point sets tailored to a specific kernel, making them sensitive to misspecification and less robust in practice. In this work, we study randomised quadrature methods with a focus on robustness rather than kernel-specific optimality. By minimising a tractable upper bound on the worst-case error, we obtain an explicit sampling distribution p*∝ πg with g=2d/(2α+d), which depends on the integration density π and on a but not on the kernel beyond its Sobolev order. Under a weak doubling condition on the design measure, independent samples from p* attain the minimax rate n-α/d. These assumptions cover a broad class of targets on compact and unbounded domains; we verify them explicitly for Beta-type densities, Gaussian measures, and Student-t distributions, the last of which yields the minimax rate n-min(α,(n+d/2)/d. This kernel-agnostic design improves robustness while maintaining optimal rates, and it applies beyond compact domains. The results provide both theoretical guarantees and a practical recipe for robust, rate-optimal randomised quadrature.

estimating evidence redux

Posted in Books, Statistics, University life with tags , , , , , , , , on November 21, 2025 by xi'an

Following our arXival on the new version of our HPD based Gelfand & Dey estimator of evidence, I got pointed at Wang et al. (2018), which I had forgotten I had read at the time (as testified by an ‘Og entry). Reading my own comments, I concur (with myself¹⁸!) that the method is not massively compelling since it requires a partition set that is strongly related with the targeted integral. The above illustration for a mixture, that is for a pseudo posterior that is a mixture with two Gaussian components with known variance, also shows (in reverse) the curse of dimension and the need for finely tuned partitions. Said partition corresponding to the myriad of sets on the rhs. With such a degree of partitioning, Riemann integration should also produce perfect estimate, as shown by the zero error in the resulting estimator (Table 4).

finite variance goals

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , on November 8, 2025 by xi'an

During Johan Seger’s seminar in Warwick, on the control variate improvements he developed with Rémi Leluc (which PhD thesis committee I joined), Aymeric Dieuleveut, François Portier, and Aigerim Zhuman, I started wondering at whether or not a control variate could turn an infinite variance Monte Carlo estimate into a finite variance one. And asked… ChatGPT about it, with the above reply that is correct if not practical in the least since the example provided therein was reverse-engineering an infinite variance rv into a sum of an infinite variance rv considered as the control variate and a finite variance rv. As summarised below. In practice, this would mean replacing the integrand of interest with a much simpler integrand that shares the same asymptotic behaviour, not an easy task! (As an aside, I found out that enabling MathJax on this ‘Og would cost me $40 a month!)

5. Summary

✅ Theoretical possibility:
Yes — control variates can make an infinite-variance estimator finite, but only if the control’s sample path shares the same tail driver and its expectation is known.infinite variance rv
In real-world Monte Carlo, when X is heavy-tailed, you usually:

  1. Split X = Y + (X-Y), where Y has known expectation and similar tails,

  2. Use Y as control variate, and

  3. Possibly combine with truncation, conditional expectation, or importance sampling for stability.

coupling-based approach to f-divergences diagnostics for MCMC

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , on October 27, 2025 by xi'an

Adrien Corenflos (University of Warwick) and Hai-Dang Dau (NUS) just arXived their paper on MCMC diagnostics that Adrien told me about last month, while in Warwick.

“This [f-divergence] bound is clearly suboptimal since it does not vary in t and does not take into account the mixing of the Markov chain. We present a scheme where the weights are ‘harmonized’ as the Markov chain progresses, reflecting its mixing through the notion of coupling.”

They start by opposing the classical ergodic average and embarrassingly parallel estimates obtained by N parallel chains culled of their B initial values, to couplings used in standard diagnoses. Opting for the parallel perspective, maybe rekindling the diagnostic war of the early 1990s! The evaluation tool in the paper is based on f-divergences, like the χ² divergence which naturally relates to the effective sample size when considering weighted atomic measures. When consistent, these weighted approximations produce upper bounds on the f-divergence, with exact convergence in case of independence.

In my opinion the most exciting part of the paper stands with the ability to modify these weights along MCMC iterations, since the naïve sequential importance sampling argument I also use in class keeps them constant! The trick is to (be able to) couple randomly chosen parallel chains, with the weights being averaged at each coupling event. The resulting algorithm preserves expectation (in the importance sampling sense) and consistency (in the particle sense). Furthermore, the f-divergence bound based on the weights can only decrease between iterations, which reminds me of interleaving. And exponential convergence of the weights to uniform ones (under the strong assumption of a uniformly lower bounded probability of coupling). The paper concludes with interesting remarks on perfect sampling, Rao-Blackwellisation, control variates, and backward sampling.

A long-standing gap exists between the theoretical analysis of Markov chain Monte Carlo convergence, which is often based on statistical divergences, and the diagnostics used in practice. We introduce the first general convergence diagnostics for Markov chain Monte Carlo based on any f χ² -divergence, allowing users to directly monitor, among others, the Kullback–Leibler and the divergences as well as the Hellinger and the total variation distances. Our first key contribution is a coupling-based ‘weight harmonization’ scheme that produces a direct, computable, and consistent weighting of interacting Markov chains with respect to their target distribution. The second key contribution is to show how such consistent weightings of empirical measures can be used to provide upper bounds to f -divergences in general. We prove that these bounds are guaranteed to tighten over time and converge to zero as the chains approach stationarity, providing a concrete diagnostic.

mostly [14] M[ar]C[h] seminar

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , on March 8, 2025 by xi'an