One morning session on optimal transport after first-hand witnessing the impressive ballet of orderly lines entering the subway at Nagoya Station (and a very early run along the river and a high humidity rate, hence the picture of empty street at 5am). With Hugo Lavenant exhibiting optimal rates for Bayesian nonparametrics using Wasserstein distances, Pierre Jacob coupling MCMC chains, and Anya Katsevich investigating non-Gaussian asymptotic distributions in high dimensions (beyond Bernstein-von Mises). Then I tried to attend the session on Bayesian Uncertainty Quantification and Posterior Sampling for Large-Scale Generative Models, but it proved too popular for the number of seats, and I ended up discussing with others. After a nap related to my jetlag induced, early, rise I went back to chair Sid Chib’s Foundation Lecture, where he discussed the use of an orbit of models in Bayesian model choice, rather than the (MAP) most likely one. Based on their 2018 JASA paper which I already discussed in Paris with Anna Simoni presenting. Hence reminding me of points I presumably already made, from the issue of having too many models to realistically explore to constructing coherent priors across them, with the fractional, empirical, proxy of using a (same) fraction of the sample as a learning sample, to more philosophical issues like missing a utility function about having to chose a model, especially with all models being wrong, missing an uncertainty quantification on the evidence itself, rather than using most likely models (MAP!), called an orbit by Sid (which requires some calibration). Since the uncertainty represented by the sample induces an uncertainty in the ranking of models. And a most appropriate, almost local, occurrence of the Rashomon principle!!! And I finished the day mixing with many friends in the poster session¹, where Darren Wraith presented our ongoing work on novel, adaptive, importance, sampling.
Archive for Bernstein-von Mises theorem
ISBA 2026²
Posted in Books, pictures, Running, Statistics, Travel, University life with tags Bayesian model averaging, Bayesian model choice, Bernstein-von Mises theorem, Chib's approximation, evidence, ISBA 2026, ISBA World Meeting, Japan, map, model uncertainty, Nagoya, Rashomon, Sid Chib, Wasserstein distance on July 1, 2026 by xi'anMasterclass in Bayesian Asymptotics, Université Paris Dauphine, 18-22 March 2024
Posted in Books, pictures, Statistics, Travel, University life with tags 2024, Bayesian asymptotics, Bayesian foundations, Bayesian inference, Bernstein-von Mises theorem, bois de Boulogne, course, empirical Bayes methods, foundation lectures, France, graduate course, IMS Lecture Notes, Judith Rousseau, marginal likelihood, MASH, Master program, masterclass, Paris, PariSanté campus, Université Paris Dauphine, University of Oxford on December 8, 2023 by xi'an
On the week of 18-22 March 2024, Judith Rousseau (Paris Dauphine & Oxford) will teach a Masterclass on Bayesian asymptotics. The masterclass takes place in Paris (on the PariSanté Campus) and consists of morning lectures and afternoon labs. Attendance is free with compulsory registration before 11 March (since the building is not accessible without prior registration).
The plan of the course is as follows
Part I: Parametric models
In this part, well- and mis-specified models will be considered.
– Asymptotic posterior distribution: asymptotic normality of the posterior, penalization induced by the prior and the Bernstein von – Mises theorem. Regular and nonregular models will be treated.
– marginal likelihood and consistency of Bayes factors/model selection approaches.
– Empirical Bayes methods: asymptotic posterior distribution for parametric empirical Bayes methods.
Part II: Nonparametric and semiparametric models
– Posterior consistency and posterior convergence rates: statistical loss functions using the theory initiated by L. Schwartz and developed by Ghosal and Van der Vaart, results on less standard or well behaved losses.
– semiparametric Bernstein von Mises theorems.
– nonparametric Bernstein von Mises theorems and Uncertainty quantification.
– Stepping away from pure Bayes approaches: generalized Bayes, one step posteriors and cut posteriors.
BayesComp²³ [aka MCMski⁶]
Posted in Books, Mountains, pictures, Running, Statistics, Travel, University life with tags 0.234, ABC, agent-based models, BayesComp 2025, Bayesian GANs, Bernstein-von Mises theorem, Hokkaido, maximum mean discrepancy, Milano, misspecification, noise contrasting estimation, posters, qMC, quadrature, Rademacher complexity, score function, simulated annealing, SMC, Steve Fienberg, Wasserstein distance on March 20, 2023 by xi'an
The main BayesComp meeting started right after the ABC workshop and went on at a grueling pace, and offered a constant conundrum as to which of the four sessions to attend, the more when trying to enjoy some outdoor activity during the lunch breaks. My overall feeling is that it went on too fast, too quickly! Here are some quick and haphazard notes from some of the talks I attended, as for instance the practical parallelisation of an SMC algorithm by Adrien Corenflos, the advances made by Giacommo Zanella on using Bayesian asymptotics to assess robustness of Gibbs samplers to the dimension of the data (although with no assessment of the ensuing time requirements), a nice session on simulated annealing, from black holes to Alps (if the wrong mountain chain for Levi), and the central role of contrastive learning à la Geyer (1994) in the GAN talks of Veronika Rockova and Éric Moulines. Victor Elvira delivered an enthusiastic talk on our massively recycled importance on-going project that we need to complete asap!
While their earlier arXived paper was on my reading list, I was quite excited by Nicolas Chopin’s (along with Mathieu Gerber) work on some quadrature stabilisation that is not QMC (but not too far either), with stratification over the unit cube (after a possible reparameterisation) requiring more evaluations, plus a sort of pulled-by-its-own-bootstrap control variate, but beating regular Monte Carlo in terms of convergence rate and practical precision (if accepting a large simulation budget from the start). A difficulty common to all (?) stratification proposals is that it does not readily applies to highly concentrated functions.
I chaired the lightning talks session, which were 3mn one-slide snapshots about some incoming posters selected by the scientific committee. While I appreciated the entry into the poster session, the more because it was quite crowded and busy, if full of interesting results, and enjoyed the slide solely made of “0.234”, I regret that not all poster presenters were not given the same opportunity (although I am unclear about which format would have permitted this) and that it did not attract more attendees as it took place in parallel with other sessions.
In a not-solely-ABC session, I appreciated Sirio Legramanti speaking on comparing different distance measures via Rademacher complexity, highlighting that some distances are not robust, incl. for instance some (all?) Wasserstein distances that are not defined for heavy tailed distributions like the Cauchy distribution. And using the mean as a summary statistic in such heavy tail settings comes as an issue, since the distance between simulated and observed means does not decrease in variance with the sample size, with the practical difficulty that the problem is hard to detect on real (misspecified) data since the true distribution behing (if any) is unknown. Would that imply that only intrinsic distances like maximum mean discrepancy or Kolmogorov-Smirnov are the only reasonable choices in misspecified settings?! While, in the ABC session, Jeremiah went back to this role of distances for generalised Bayesian inference, replacing likelihood by scoring rule, and requirement for Monte Carlo approximation (but is approximating an approximation that a terrible thing?!). I also discussed briefly with Alejandra Avalos on her use of pseudo-likelihoods in Ising models, which, while not the original model, is nonetheless a model and therefore to taken as such rather than as approximation.
I also enjoyed Gregor Kastner’s work on Bayesian prediction for a city (Milano) planning agent-based model relying on cell phone activities, which reminded me at a superficial level of a similar exploitation of cell usage in an attraction park in Singapore Steve Fienberg told me about during his last sabbatical in Paris.
In conclusion, an exciting meeting that should have stretched a whole week (or taken place in a less congenial environment!). The call for organising BayesComp 2025 is still open, by the way.
BNP13
Posted in Mountains, pictures, Running, Statistics, Travel with tags Bayesian non-parametrics, Bernstein-von Mises theorem, Biometrika, BNP13, Bruno de Finetti, Charles de Gaulle, Chile, conference, ISBA, jetlag, label switching, Lago Llanquihue, optimal coupling, optimal transport, parallel MCMC, Patagonia, Puerto Varas on October 28, 2022 by xi'an
BNP13 is set in this incredible location on a massive lake (almost as large as Lac Saint Jean!) facing several tantalizing snow-capped volcanoes… My trip from Paris to Puerto Varas was quite smooth if relatively longish (but I slept close to 8 hours on the first leg and busied myself with Biometrika submissions the rest of the way). Leaving from Paris at midnight proved a double advantage as this was one of the last flights leaving, with hardly anyone in the airport. On Sunday, I arrived early enough to take a quick dip in Lake Llanquihue which was fairly cold and choppy!
Overall the conference is quite exhilarating as all talks are of interest and often covering on-going research. This may be one of the most engaging meetings I have attended in the past years! Plus a refreshing variety of topics and seniority in the speakers.
To start with a bang!, Sonia Petrone (Bocconi) gave a very nice plenary lecture in the most auspicious manner, covering her recent works on Bayesian prediction as an alternative way to run Bayesian inference (in connection with the incoming Read Paper by Fong et al.). She covered so much ground that I got lost before long (jetlag did not help!). However, an interesting feature underlying her talk is that, under exchangeability, the sequence of predictives converges to a random probability measure, a de Finetti way to construct the prior that is based on predictives. Avoiding in a sense the model and the prior on the parameters of that process. (The parameter is derived from the infinite exchangeable [or conditionally iid] sequence, but the sequence of predictives need be defined.) The drawback is that this approach involves infinite sequences, with practical truncation to a finite horizon being an approximation whose precision / error may prove elusive to characterise. The predictive approach also allows to recover a limiting Normal distribution (not a Bernstein-von Mises type!) and hence credible intervals on parameters and distributions.
While this is indeed a BNP conference (!), I was surprised to see lot of talks paying attention to clustering and even to mixtures, with again a recurrent imprecision on the meaning of a cluster. (Maybe this was already the case for BNP11 in Paris but I may have been too busy helping with catering to notice!) For instance, Brian Trippe (MIT) gave a quick intro on his (AISTATS 2022) work on parallel MCMC with coupling. As unbiased MCMC strongly improving upon naïve parallel MCMC relative to the computing cost. With an interesting example where coupling is agnostic to the labeling of random partitions in clustering problems, involving optimal transport, manageable in O(K³log(K)) time when K is the number of clusters.
EM degeneracy
Posted in pictures, Statistics, Travel, University life with tags ABC, BayesComp 2020, Bernstein-von Mises theorem, clustering, compatible conditional distributions, conference, cut models, cycle path, EM algorithm, Gibbs sampling, hidden Markov models, Institut de Mathématique d'Orsay, MCMC, MHC 2021, mixtures, particle filters, physical attendance, Rao-Blackwellisation, SEM, SMC, smoothing, Université Paris-Sud on June 16, 2021 by xi'an
At the MHC 2021 conference today (to which I biked to attend for real!, first time since BayesComp!) I listened to Christophe Biernacki exposing the dangers of EM applied to mixtures in the presence of missing data, namely that the algorithm has a rising probability to reach a degenerate solution, namely a single observation component. Rising in the proportion of missing data. This is not hugely surprising as there is a real (global) mode at this solution. If one observation components are prohibited, they should not be accepted in the EM update. Just as in Bayesian analyses with improper priors, the likelihood should bar single or double observations components… Which of course makes EM harder to implement. Or not?! MCEM, SEM and Gibbs are obviously straightforward to modify in this case.
Judith Rousseau also gave a fascinating talk on the properties of non-parametric mixtures, from a surprisingly light set of conditions for identifiability to posterior consistency . With an interesting use of several priors simultaneously that is a particular case of the cut models. Namely a correct joint distribution that cannot be a posterior, although this does not impact simulation issues. And a nice trick turning a hidden Markov chain into a fully finite hidden Markov chain as it is sufficient to recover a Bernstein von Mises asymptotic. If inefficient. Sylvain LeCorff presented a pseudo-marginal sequential sampler for smoothing, when the transition densities are replaced by unbiased estimators. With connection with approximate Bayesian computation smoothing. This proves harder than I first imagined because of the backward-sampling operations…