A few days before the January OWABI, I read through Simon Kucharsky’s and Paul Bürkner’s paper, arXived on 17 January. Which proposed an amortized Bayesian inference (ABI) method, even though the ABI is not the same as in OWABI! The motivation for their work is to start from a (standard) mixture model where the components are not analytically tractable (but still parameterised). But a generative model nonetheless. As in the earlier reviewed paper (which was arXived on the same day), by MEJ Newman, the dual representation of the joint posterior p(θ,z|x) as p(z|x,θ)p(θ|x) and p(θ|z,x)p(z|x) is (over?) emphasized (albeit unclearly why!). ABI uses neural networks and more specifically normalising flows to approximate the posterior p(θ|x) from prior predictive samples (θ,x) (as in ABC), and then directly exploit the invertibility of said flows to generate from this approximate posterior. One interesting aspect of the modelling is the derivation of summary statistics in the design of the network, albeit mixture posteriors do not allow for dimension-reduced (Bayes) sufficient statistics (and a contradictory sentence that conditioning on the summaries “does not alter the target posterior”, p7). The resulting approximate posterior generator proves much much faster than running an MCMC, obviously, and furthermore adapt to handling a sequence of datasets. A second network is constructed to approximate p(z|x,θ), using the same summaries. The network parameters are estimated through losses, rather than in a Bayesian manner, with a default Kullback-Leibler version (18). I also fail to understand why the networks are trained over unconstrained parameters when all parameters could become unconstrained when using the adequate parameterisation. And am fairly surprised at the regression towards the ill-fated step of using ordered parameters to avoid label switching… But the main quandary remains the issue of assessing the approximation effect, despite experiments aiming at pacifying such worries. And similarities with Stan and BayesFlow.
Archive for summary statistics
amortized Bayesian mixture model
Posted in Books, Statistics, University life with tags ABC, amortization, amortized Bayesian inference, Approximate Bayesian computation, approximate Bayesian inference, dimension reduction, doubly intractable posterior, label switching, mixtures of distributions, normalising flow, OWABI, reparameterisation, STAN, sufficiency, summary statistics, webinar on February 7, 2025 by xi'anveniSBA²
Posted in Books, pictures, Running, Statistics, Travel, University life with tags ABC, Bayesian conference, Bayesian data analysis, Bayesian decision theory, BDA, canal, coherence, composite likelihood, ISBA 2024, Italia, mixtures of experts, neural network, normalising flow, One World ABC Seminar, prior selection, score function, sequential Monte Carlo, simulation-based inference, SIS model, slaughterhouse, subjective prior, summary statistics, Università Ca' Foscari Venezia, Venezia, Venice on July 4, 2024 by xi'an
After another morning cycle of 2Xing Porte della Libertà (under a light and pleasant rain) and swimming in Sant’ Alviso (in too warm a water), I did not make it for the beginning of the Bayesian deep learning session, breakfast oblige!, and cumulated with different percolation events (ie, meeting friend after friend on my way to the classroom), I could not get enough of the session to report anything even barely useful!
As I did not rush fast enough to Andrew’s Foundation lecture (another sequence of percolations!), I had to stand in the back of the packed main amphitheatre (and former sorting hall of the Venice slaughterhouse!), Guido Cazzavillan’s Aula Magna, while he talked a fresco about some holes in Bayesian data analysis (the analysis, not the book!), those being [verbatim]
- the usual rules of conditional probability fail in the quantum realm,
- flat or weak priors lead to terrible inferences about things we care about,
- subjective priors are incoherent,
- Bayesian decision picks the wrong model,
- Bayes factors fail in the presence of flat or weak priors,
- for Cantorian reasons we need to check our models, but this destroys the coherence of Bayesian inference.
After lunch, I attended the (mostly sequential) simulation based inference (renamed from ABC!) session with a composite likelihood proposal by Lorenzo Rimella, that uses marginals to approximate the likelihood of a hidden Markov SIS epidemic model by composite likelihood towards getting more efficient if inexact versions. Then [1WABC webinar co-organiser] Umberto Picchini on surrogates for likelihood and posterior functions, with sequential improvements (w/o ABC and w/o neural networks). Called “Sequential mixture posterior and likelihood estimation”, using mixtures of experts when the weights are functions of the observed or simulated y. With adapting the number of components in the mixture. Comparing favourably with normalising flows. And Wentao Li on correcting by ABC for composite likelihood as in Ruli et al. (2016). Where a posterior distribution given composite scores (seen as [summary] statistics) is employed but requires a convergent estimator of the unknown parameter.
No congratulation today to our PhD student who managed to fall in a canal (but survived)..!

Mark [a]B[c], plus cats
Posted in Books, Kids, pictures, Statistics, University life with tags ABC, ABC variable selection, Bayesian neural networks, HDR, HPD region, Mark Beaumont, model misspecification, normalizing flow, One World ABC Seminar, philogenic trees, principal components, prior predictive, Saving Wildcats, Scotland, summary statistics, The Guardian, University of Bristol, University of Warwick, webinar, wildcats on June 22, 2024 by xi'an
30 May was a day of first and last times, if not in capital ways (for me), on the One World ABC webinar. This was the first time we had a talk by Mark Beaumont and also the first time I had team experience of facing a smoking participant, while this was the last time of our monthly webinar for the (Northern) academic year.
The talk was about model misspecification in population genomic, from an ABC perspective with the motivation of common noticeable difference between the distributions of d(s,s⁰) and d(s,s’), distances between the prior predictively simulated summary statistics and the observed ones vs posterior generated ones, which should indicates misspecification, esp with complicated models. Mark and his coauthors then supported a gradual elimination of summary statistics to diminish the discrepancy, hence voluntarily impoverishing the model. While blaming the statistics sounded a bit like shooting the messenger, the resolution is of obvious interest if backing from modelling the misspecification itself.
Some of the presented work was conducted in Ward et al (2022, NeurIPS) with a reference to the outlying Ratmann et al (2009) we later discussed, for including tolerance as an extra parameter ε, thereby reconsidering Wikinson’s exact ABC for noisy observations y by the medium of a normalising flow on the marginal distribution of the denoised x (learned from the prior predictive)
More precisely, the idea is to drop summary statistics by checking whether or not the observed S⁰ belongs to HPD region, removing one component of S at a time, using e.g. a k-NN estimate for the summary density (hence depending on parameterisation of said statistics for the distance). Hopefully, the process stops before loosing identifiability by using too few statistics. I also wondered at multiple uses of the data in this sequential procedure but Mark argued for adopting a meta- or pragma- or Gelmanian- Bayesian perspective in the end!
Another perk was the appearance of (and illustration with) the Scottish Wildcat, mentioned in The Guardian a few months ago and discussed in the ‘Og, with further papers exploring more aspects of this hybridization, like a posterior applied to a much more complex phylogenic tree reconstruction for cats of different creeds and many related parameters.
Marc Beaumont on One World ABC webinar [30 May, 9am]
Posted in Books, pictures, Statistics, Travel, University life with tags ABC, ABC consistency, ABC-PMC, computational statistics, likelihood-free methods, Mark Beaumont, misspecification, One World ABC Seminar, population genomics, population Monte Carlo, simulation, summary statistics on May 17, 2024 by xi'an
For the final talk of this Spring season of the One World ABC webinar, we are very glad to welcome Marc Beaumont, a central figure in the development of ABC methods and inference! (And a coauthor of our ABC-PMC paper.)
Model misspecification in population genomic
Mark Beaumont
University of Bristol
30th May 2024, 9.00am UK time
Abstract
In likelihood-free settings, problematic effects of model misspecification can manifest themselves during computation, leading to nonsensical answers, particularly causing convergence problems in sequential algorithms. This issue has been well studied in the last 10 years, leading to a number of methods for robust inference. In practical applications, likelihood-free methods tend to be applied to the output of complex simulations where there is a choice of summary statistics that can be computed. One approach to handling misspecification is to simply not use summary statistics computed from simulations of the model under the prior that cannot be with those observed in the data. This presentation gives a brief review of methods for observing and handling misspecification in ABC and SBI, and then discusses approaches that we have explored in a population genomic modelling framework.
Asymptotics of ABC when summaries converge at heterogeneous rates
Posted in pictures, Statistics, University life with tags ABC, Approximate Bayesian computation, Bayesian consistency, COVID-19, curse of dimensionality, lockdown, PhD thesis, summary statistics, Université Paris Dauphine, University of Oxford on November 21, 2023 by xi'an
We just posted a new arXival, jointly with Caroline Lawless, Judith Rousseau, and Robin Ryder. This is a significant component of Caroline’s PhD thesis in Oxford, on which we started working during the first COVID lockdown. In this paper, we extend our results with David Frazier, Gael Martin, both with whom I’ll soon be reunited!, and Judith, published in Biometrika in 2018, to the more challenging case where different components of the summary statistic vector converge to their respective means at different rates, with some possibly not even converging at all. While this sounds impossible (!), we do prove consistency of the ABC posterior under such heterogeneous rates.
Wentao Li and Paul Fearnhead (also in Biometrika and in 2018) reduce the curse of the dimension of the set of summary statistic by showing, in the specific case of asymptotically normal summary statistics concentrating at the same rate, that a local linear post-processing step leads to a significant improvement in the theoretical behaviour of the ABC posterior. However, due to this focus on reducing the impact of the dimension of the summary statistics, it is therefore important to study its efficiency in a context where the summary statistics are not as well behaved. Surprinsingly maybe, we show that the significant improvement due to local linear post-processing persists even when summary statistics have heterogeneous behaviour. Most interestingly, the number of summary statistics which converge at the fast rate has no impact on the rate of posterior concentration nor on the shape of the ABC posterior (provided it exceeds the dimension of the parameter).