Archive for self-normalised importance sampling

mostly Monte Carlo [09/10, PSC]

Posted in Statistics, University life with tags , , , , , , , , , , , , on October 4, 2026 by xi'an

The next episode of our mostly Monte Carlo seminar is next Friday (9 October) at PariSanté Campus (room #8) with speakers

15:00 – Víctor Elvira, University of Edinburgh

16:00 – Edoardo Bandoni, Université Paris Dauphine-PSL

Víctor Elvira, “Rethinking self-normalized importance sampling”

Self-normalized importance sampling (SNIS) is one of the most widely used Monte Carlo techniques for inference with unnormalized target distributions. Despite its usefulness, SNIS is often viewed simply as a normalized version of ordinary importance sampling, and many of its methodological questions remain largely unexplored. In this talk, we revisit SNIS from a unified perspective. We first introduce a generalized formulation of self-normalized importance sampling based on coupled proposals, showing that the classical SNIS estimator is only one member of a broader family of Monte Carlo estimators with new opportunities for variance reduction. We then consider the classical SNIS estimator and present adaptive algorithms that learn proposals tailored to its optimal proposal distribution, together with theoretical guarantees including consistency, asymptotic normality, and convergence of the proposal. Together, these developments suggest that self-normalized importance sampling should be regarded as a distinct Monte Carlo methodology, with its own theory, optimality principles, and algorithmic design.

Edoardo Bandoni, “Rate-Optimal Randomised Kernel Quadrature”

Kernel quadrature is widely used to approximate integrals of smooth functions, with the worst-case error typically decaying at the minimax rate n-α/d for smoothness α in dimension d. Existing rate-optimal methods often depend on deterministic point sets tailored to a specific kernel, making them sensitive to misspecification and less robust in practice. In this work, we study randomised quadrature methods with a focus on robustness rather than kernel-specific optimality. By minimising a tractable upper bound on the worst-case error, we obtain an explicit sampling distribution p*∝ πg with g=2d/(2α+d), which depends on the integration density π and on a but not on the kernel beyond its Sobolev order. Under a weak doubling condition on the design measure, independent samples from p* attain the minimax rate n-α/d. These assumptions cover a broad class of targets on compact and unbounded domains; we verify them explicitly for Beta-type densities, Gaussian measures, and Student-t distributions, the last of which yields the minimax rate n-min(α,(n+d/2)/d. This kernel-agnostic design improves robustness while maintaining optimal rates, and it applies beyond compact domains. The results provide both theoretical guarantees and a practical recipe for robust, rate-optimal randomised quadrature.

inverse probability weighting

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , on May 4, 2026 by xi'an

Quite recently, Jyotishka Datta and Nick Polson published a fairly interesting [imho] paper in The New England Journal of Statistics in Data Science, entitled Inverse Probability Weighting: From Survey Sampling to Evidence Estimation that (obviously) caters to my own interests! They bring three threads together. First, they recall the long debate between using [normalized] Horvitz–Thompson and [self-normalized] Hájek estimators in survey sampling, pointing out that the latter is “usually the better estimator, despite estimation of an a priori known quantity” (citing from Särndal & al., 2003). Which is also my experience with importance sampling, as in this 1995 Note aux Comptes Rendus with George. The mathematical paradox of “estimating” a constant is central to other advances in the area, like noise-contrastive estimation à la Gutmann & Hyvärinen (2005) or the measure estimation of Kong & al. (2003). Even more interestingly, Datta & Polson consider there is a link with the inconsistent Bayesian (counter)example of Larry Wasserman and Jamie Robins, where the censoring probability increases with the value of the parameter of interest, paradox in which Chris Sims also got involved. (I was unaware that he had passed away last month.) And, lo and behold!, with the Stein “paradox” of my PhD years (and beyond).

Their central argument stands with the missing data link between Horvitz-Thompson survey sampling and Monte Carlo integration. (With a reference to our Riemann sum papers with Anne Philippe!, making me realise the authors had recently published an extension on that idea.) This reminds me very much of the missing measure approach of Kong et al. (2003). (Actually the reference appears in the final discussion.) The authors go over several paradoxes like Basu’s circus estimate (1988), Larry’s inconsistent Bayes estimate (2004), where Horvitz–Thompson performs nicely under compactness assumptions, the Bayesian answers (which include nested sampling even though I do not see the connection). Especially Li’s (2010) solution.  The attached numerical experiment displays a consistent underperformance of the Horvitz-Thompson estimator, in contrast with the theory… 

In conclusion, while enjoying very much revisiting so many examples and papers I came across in the past decades, I remain somewhat puzzled by the lack of overall message.

 

BayesComp 2025.2

Posted in Kids, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on June 19, 2025 by xi'an


The main BayesComp²⁵ conference started with Pierre Jacob’s plenary talk on his recent advances on coupling for unbiased MCMC—currently ERC grantee on that topic—.  Raising lazy questions like using a different target or transition kernel for the second chain in the coupling, connecting the Poisson equation and control variates, handling the signed issue with the unbiased approximations. Interestingly, they obtain an unbiased estimator of the asymptotic variance of the unbiased estimator. And a correction for self-normalised importance sampling, which has some connections with our 1996 (?) pinball sampler. Also an evaluation of the median of means, rather than the average of means, which is a thing I had been (lazily) contemplating for a while  (On the greedy side, as I was writing my recovery exam for my Monte Carlo course, I realised the results Pierre presented could be somewhat recycled into exam problems!)

My first parallel session was on gradient-based methods with a talk by Francesca Crucinio on proximal particle Langevin algorithms (similar to the one she gave in PariSanté last year), a talk by Zhihao Wang on stereographic multiple try Metropolis(-Rosenbluth-Teller) that unsurprisingly recovers ergodicity thanks to the compactness of the ball. For which I wonder why a Normal proposal makes complete sense since one could consider a mover after the projection instead and why iid rather than repelling multiple proposals are used… The last speaker was just out from the plane from California, Siddharth Vishwanath who spoke about repelling-attracting HMC. With very nice animations of HMC, if reaching the main point of using both negative and positive frictions a few minutes before the session finished. The method preserves volume and potential, if not energy.

Speaking of which (energy), I find myself struggling with my less than 6 hours of sleep since arrival during the first afternoon session, despite a fiery hot spot lunch, which means in plainer terms that I alas dozed in and out of the talks. The second session saw Jack Jewson exposing in deeper details the exact PDMP algorithm for Gibbs measures  Jeremias Knoblauch mentioned yesterday. And Jonathan Huggins as well, using Gaussian processes as proxies for expected likelihoods, with lower guarantees than pseudo-marginal versions. In a mildly connected way, Robin Ryder went through the resolution of the ecological inference challenge they produce with Nicolas Chopin and Théo Valdoire (all authors with whom I am connected, Théo being a brillant student of our MASH Master last year and now in Harvard, hopefully till the end of his PhD!)

On the extra-academic curriculum, I had a yummy dinner in the Maxwell Hawker (street) food centre, incl. Xiao Long Bao that cooled down fast enough to avoid the usual scalding effect, plus rojak a mixed fruit and vegetable fried in a peanut sauce that I had never tasted before, popiah (ditto), chili noodles, and an appam with durian deepfried balls as a fabulous and unexpected dessert.

exceptional OWABI web/sem’inar [19 June, BayesComp²⁵]

Posted in pictures, Statistics, Travel, Uncategorized, University life with tags , , , , , , , , , , , , , , on June 10, 2025 by xi'an


Exceptionally, the next One World Approximate Bayesian Inference (OWABI) Seminar will be hybrid as it is scheduled to take place during BayesComp 2025 in Singapore, on Thursday 19 June at 8pm Singapore time (1pm in Tórshavn) and two talks, one by Filippo Pagani on

Approximate Bayesian Fusion
Bayesian Fusion is a powerful approach that enables distributed inference while maintaining exactness. However, the approach is computationally expensive. In this work, we propose a novel method that incorporates numerical approximations to alleviate the most computationally expensive steps, thereby achieving substantial reductions in runtime. Our approach retains the flexibility to approximate the target posterior distribution to an arbitrary degree of accuracy, and is scalable with respect to both the size of the dataset and the number of computational cores. Our method offers a practical and efficient alternative for large-scale Bayesian inference in distributed environments.
and one by Maurizio Filippone on
GANs Secretly Perform Approximate Bayesian Model Selection
Generative Adversarial Networks (GANs) are popular models achieving impressive performance in various generative modeling tasks. In this work, we aim at explaining the undeniable success of GANs by interpreting them as probabilistic generative models. In this view, GANs transform a distribution over latent variables Z into a distribution over inputs X through a function parameterized by a neural network, which is usually referred to as the generator. This probabilistic interpretation enables us to cast the GAN adversarial-style optimization as a proxy for marginal likelihood optimization. More specifically, it is possible to show that marginal likelihood maximization with respect to model parameters is equivalent to the minimization of the Kullback-Leibler (KL) divergence between the true data generating distribution and the one modeled by the GAN. By replacing the KL divergence with other divergences and integral probability metrics we obtain popular variants of GANs such as f-GANs, Wasserstein-GANs, and Maximum Mean Discrepancy (MMD)-GANs. This connection has profound implications because of the desirable properties associated with marginal likelihood optimization, such as (i) lack of overfitting, which explains the success of GANs, and (ii) allowing for model selection, which opens to the possibility of obtaining parsimonious generators through architecture search.

These talks will be delivered on-site and on-line, as a Zoom visio-conference.

next OWABI webinar [24 April]

Posted in pictures, Statistics, Uncategorized, University life with tags , , , , , , , , , , , , on April 16, 2025 by xi'an


The next One World Approximate Bayesian Inference (OWABI) Seminar is scheduled on Thursday the 24th of April at 11am UK time (12am CET) with the speaker being Ayush Bharti (Aalto University), who will talk about

“Cost-aware simulation-based inference “

Abstract: Simulation-based inference (SBI) is the preferred framework for estimating parameters of intractable models in science and engineering. A significant challenge in this context is the large computational cost of simulating data from complex models, and the fact that this cost often depends on parameter values. We therefore propose cost-aware SBI methods which can significantly reduce the cost of existing sampling-based SBI methods, such as neural SBI and approximate Bayesian computation. This is achieved through a combination of rejection and self-normalised importance sampling, which significantly reduces the number of expensive simulations needed. Our approach is studied extensively on models from epidemiology to telecommunications engineering, where we obtain significant reductions in the overall cost of inference. .