Archive for adaptive importance sampling

multimodal challenges

Posted in Books, Statistics, University life with tags , , , , , , , , , , on April 30, 2026 by xi'an

At the last mostly Monte Carlo seminar, Pierre Monmarché presented a recent work on post-sampling for multimodal targets: while I  consider the main problem in sampling from generic multimodal targets stands with finding the modes, rather than with exploring local aspects or estimating relative weights of said modes, this made me ponder whether or not this could be accelerated by removing chunks of the already explored modes to induce moves elsewhere, which is a form of radical, brute-force, tempering, or of Wang-Landau.  As for the relative weights, a multiple move proposal can be considered, including our folding idea. Or Geyer’s inverse logistic trick. Or the similar mixture trick we used in our Biometrika paper on nested sampling. Pierre’s approach was closer to adaptive importance sampling, with a self-imposed constraint of fixed sample sizes from (approximate) distributions around each of the modes.

Karim Benabed, astrophysician

Posted in Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on December 5, 2025 by xi'an

The astronomer and cosmologist Karim Benabed got killed on Wednesday in Paris. While cycling, run over by a truck-driver (with no further details at the moment). He was a senior research at IAP (Institut d’Astrophysique de Paris) and we actively collaborated together between 2005 and 2010 on an ANR project on efficient simulation methods for inferring cosmological parameters, based on PMC. And Bayesian model comparison. He was a very congenial person, very sharp in assimilating new methods and keen on exploring novel hypotheses. While we did not keep closely in touch, I would meet him now and then while visiting the IAP. Ironically, Darren Wraith, formerly a postdoc with him at IAP,  was visiting me last week and we were reminiscing of that era as late as Saturday night over dinner… So sad (and also so absurd, a truck stopping the trajectory of someone managing to travel to the origins of the Universe). The above is a cartoon of him drawn during his cosmic microwave background presentation during the Nuit de l’Astronomie.

optimal importance sampling for stochastic optimisation

Posted in Books, Statistics, University life with tags , , , , , , on June 6, 2025 by xi'an

A recent arXival by Liviu Aolaritei, Bart Van Parys, Henry Lam, and Michael Jordan (a co-PI in our ERC Synergy Ocean project) discusses optimal importance sampling schemes for stochastic optimisation, processed by an iterative Robbins-Munro algorithm improvement (with the Polyak-Ruppert improvement).

“Despite its popularity, IS is often described as a `double-edged sword.’ Its performance depends critically on the choice of the proposal distribution, which is typically sensitive to the underlying model”

I had never thought of optimising the importance function in this context, even after the seminaire of Tom Guédon mentioned in a recent ‘og.  The paper considers the optimisation of f(θ)=E[F(θ,X)] whose expectation is under a certain distribution P, through

“an iterative gradient-based algorithm that jointly updates the decision variable and the IS distribution without requiring time-scale separation between the two. Our method achieves the lowest possible asymptotic variance and guarantees global convergence under convexity of the objective and mild assumptions on the IS distribution family. Furthermore, we show that these properties are preserved under linear constraints by incorporating a recent variant of Nesterov’s dual averaging method”

with the Robbins-Munro algorithm updating the value of both θ and the parameter of the importance function. One specific difficulty is that the ideal importance function depends on the argument of the optimisation problem, obviously unknown, a “curse of circularity” I had not met previously. The linear constraint creates another kind of difficulty in order to derive the active constraint set while requiring an adapted (eg, absolutely continuous) importance function. This leads the authors to introduce a different (secondary) importance distribution on X, opening a Pandora box of infinite tuning that they opt to terminate by fixing an importance function at some finite stage. I am however uncertain as to how the combinatoric difficulty of exploring all active constraint sets at each iteration is handled. The paper being mostly theoretical, there is no illustration therein. Nor a computational cost evaluation.

[more than] everything you always wanted to know about marginal likelihood

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , on February 10, 2022 by xi'an

Earlier this year, F. Llorente, L. Martino, D. Delgado, and J. Lopez-Santiago have arXived an updated version of their massive survey on marginal likelihood computation. Which I can only warmly recommend to anyone interested in the matter! Or looking for a base camp to initiate a graduate project. They break the methods into four families

  1. Deterministic approximations (e.g., Laplace approximations)
  2. Methods based on density estimation (e.g., Chib’s method, aka the candidate’s formula)
  3. Importance sampling, including sequential Monte Carlo, with a subsection connecting with MCMC
  4. Vertical representations (mostly, nested sampling)

Besides sheer computation, the survey also broaches upon issues like improper priors and alternatives to Bayes factors. The parts I would have done in more details are reversible jump MCMC and the long-lasting impact of Geyer’s reverse logistic regression (with the noise contrasting extension), even though the link with bridge sampling is briefly mentioned there. There is even a table reporting on the coverage of earlier surveys. Of course, the following postnote of the manuscript

The Christian Robert’s blog deserves a special mention , since Professor C. Robert has devoted several entries of his blog with very interesting comments regarding the marginal likelihood estimation and related topics.

does not in the least make me less objective! Some of the final recommendations

  • use of Naive Monte Carlo [simulate from the prior] should be always considered [assuming a proper prior!]
  • a multiple-try method is a good choice within the MCMC schemes
  • optimal umbrella sampling estimator is difficult and costly to implement , so its best performance may not be achieved in practice
  • adaptive importance sampling uses the posterior samples to build a suitable normalized proposal, so it benefits from localizing samples in regions of high posterior probability while preserving the properties of standard importance sampling
  • Chib’s method is a good alternative, that provide very good performances [but is not always available]
  • the success [of nested sampling] in the literature is surprising.

ABC webinar, first!

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , on April 13, 2020 by xi'an

Screenshot_20200409_122723

The première of the ABC World Seminar last Thursday was most successful! It took place at the scheduled time, with no technical interruption and allowed 130⁺ participants from most of the World [sorry, West Coast friends!] to listen to the first speaker, Dennis Prangle,  presenting normalising flows and distilled importance sampling. And to answer questions. As I had already commented on the earlier version of his paper, I will not reproduce them here. In short, I remain uncertain, albeit not skeptical, about the notions of normalising flows and variational encoders for estimating densities, when perceived as a non-parametric estimator due to the large number of parameters it involves and wonder at the availability of convergence rates. Incidentally, I had forgotten at the remarkable link between KL distance & importance sampling variability. Adding to the to-read list Müller et al. (2018) on neural importance sampling.

Screenshot_20200409_124707