Archive for summary statistics
Approximate Bayesian Computation with Statistical Distances for Model Selection [OWABI, 27 Nov]
Posted in Books, Statistics, University life with tags ABC model selection, Approximate Bayesian computation, approximate Bayesian inference, Bayesian inference, curse of dimensionality, information loss, intractable likelihood, One World Approximate Bayesian Inference Seminar, OWABI, simulation, simulation-based inference, summary statistics, toad, University of Warwick, webinar on November 17, 2025 by xi'an
The next OWABI seminar is delivered by Clara Grazian (University of Sidney), who will talk about “Approximate Bayesian Computation with Statistical Distances for Model Selection” on Thursday 27 November at 11am UK time:
Abstract: Model selection is a key task in statistics, playing a critical role across various scientific disciplines. While no model can fully capture the complexities of a real-world data-generating process, identifying the model that best approximates it can provide valuable insights. Bayesian statistics offers a flexible framework for model selection by updating prior beliefs as new data becomes available, allowing for ongoing refinement of candidate models. This is typically achieved by calculating posterior probabilities, which quantify the support for each model given the observed data. However, in cases where likelihood functions are intractable, exact computation of these posterior probabilities becomes infeasible. Approximate Bayesian computation (ABC) has emerged as a likelihood-free method and it is traditionally used with summary statistics to reduce data dimensionality, however this often results in information loss difficult to quantify, particularly in model selection contexts. Recent advancements propose the use of full data approaches based on statistical distances, offering a promising alternative that bypasses the need for handcrafted summary statistics and can yield posterior approximations that more closely reflect the true posterior under suitable conditions. Despite these developments, full data ABC approaches have not yet been widely applied to model selection problems. This paper seeks to address this gap by investigating the performance of ABC with statistical distances in model selection. Through simulation studies and an application to toad movement models, this work explores whether full data approaches can overcome the limitations of summary statistic-based ABC for model choice.
Keywords: model choice, distance metrics, full data approaches
Reference: C. Grazian, Approximate Bayesian Computation with Statistical Distances for Model Selection, preprint at ArXiv:2410.21603, 2025
BayesComp 2025.4
Posted in pictures, Running, Statistics, Travel, University life with tags ABC, ABC model selection, Adam, approximate Bayesian inference, BayesComp 2025, Bayesian GANs, Bayesian lasso, Bayesian neural networks, Bayesian optimisation, Bayesian paradigm, Bayesian predictive, Bayesian robustness, Bayesian semi-parametrics, Baysian learning, BIC, chili crab, differential privacy, harmonic mean estimator, homomorphic encryption, hot pot, Laplace approximation, Les Houches, maximum mean discrepancy, mee siam, mixture estimation, National University Singapore, NUS, Peranakan cuisine, plenary speaker, power posterior, privacy laws, random kernel MCMC, RATP, RER, Roberta, safe Bayes, sequential importance sampling, shrinkage, shrinkage estimation, simulation-based inference, Singapore, SNCF, splines, Stein divergence, stochastic gradient MCMC, stochastic optimisation, summary statistics, Swendsen-Wang algorithm, Sylvia Frühwirth-Schnatter, Szechuan cuisine, treadmill, University of Warwick, unknown number of components, variational Bayes methods, Wasserstein distance, William Strawderman, WU Wirtschaftsuniversität Wien, zigzag algorithm on June 21, 2025 by xi'an
The third and final day of the (main) conference started tih Emtiyaz Khan’s plenary talk on adaptive Bayesian intelligence. Or, imho, [adaptive [Bayesian]] intelligence, with the brackets indicating redundancy since intelligence need include adaptivity and [intelligent] adaptivity need proceed in a Bayesian way! Focussing first on the Bayesian learning rule via variational Bayes (with a stress on Kingma’s 1994 Adam optimisation algorithm, the “most cited paper” [in machine learning]) where learning boils down to gradient steps (due to the exponential family structure), themselves versions of Taylor (or Laplace) approximations). With an interesting vision of Bayesian updating as accounting for prediction mismatch. (I missed the connection Roberta in IMDb appearing in one slide!)
The following session offered no dilemma [sorry, Alex, Axel, Chris, Robert, Sumeet, Victor!] since it included the federated learning session I organised, with Louis Asslet, Conor Hassan, and Jean-Michel Marin as speakers. Louis’ talk was on confidential [homomorphic] accept-reject algorithms to learn from other sources, while preserving (differential?) privacy, part of which came during Les Houches workshops I organised this Spring and the one before. Exploiting the additive features of log-likelihoods and exponential variates and adopting a testing perspective on privacy. Conor motivated his model with the Australian cancer atlas project Kerrie Mengersen and others have been developing over the years. The federated approach relies on variational approximations that return the same answer as an exact resolution, but more efficiently. (From a privacy perspective, I wonder at the impact of variational approximations on protecting the data, which boils down to a choice of (sufficient) statistics for the exponential families behind those approximations.) For more complicated models incorporating spatial dependence prohibits full Bayesian inference, unfortunately. Jean-Michel commented on the richness of methods for simulation-based inference, incl. model choice. His focus was on using sequential neural likelihood estimation and sequential importance sampling to approximate evidence. As in the Read Paper of Del Moral et al. (2006). Mentioning a neural version of the harmonic mean estimator by Spurio Mancini et al. (2023)! I wondered at the degree of (Rao-Blackwell) recycling involved in the computation, Jean-Michel’s answer being that AMIS is soon coming [in a theatre near you!].

The afternoon sessions did offer any reprieve in the choice of topic! I first went to Approximate Methods for Accelerated Sampling, with Rong Tang evaluating the informativeness of summary statistics through a divergence evaluation. Using autoencoders to replace the intractable posterior, with sliced minimal model discrepancy (MMD) and (pseudo?) score matching loss for divergences (reminding me of indirect inference and synthetic likelihood). Yun Yang discussed a variational proposal to estimate the number of components in a mixture model. Surprising given the multimodal structure of mixture posteriors. And the overall irregularity of (evil!) mixture models. But I could not figure out from the talk the form of the approximation.

On the food scene, tasted a nice and spicy Peranakan rice vermicelli dish called Mee Siam yesterday in a campus restaurant, which sustained me fore the rest of the day, including the ABC s/webinar. And another spicy hot pot today at NUS, to catch up on veggies, while missing the chili crab local specialty on that trip.
next OWABI webinar [27 March]
Posted in pictures, Statistics, Uncategorized, University life with tags Approximate Bayesian computation, approximate Bayesian inference, Bayesian deep learning, Bayesian inference, conformal prediction, convolutional neural networks, dropout, intractable likelihood, neural posterior estimation, One World ABC Seminar, One World Approximate Bayesian Inference Seminar, OWABI, simulation-based inference, summary statistics, Université de Montpellier, University of Warwick, webinar on March 25, 2025 by xi'an
The next One World Approximate Bayesian Inference (OWABI) Seminar is scheduled on Thursday the 27th of March at 11am UK time (12am CET) with the speaker being Meïli Baragatti (Université de Montpellier)
Approximate Bayesian Computation with Deep Learning and Conformal Prediction
Abstract: Approximate Bayesian Computation (ABC) methods are commonly used to approximate posterior distributions in models with unknown or computationally intractable likelihoods. Classical ABC methods are based on nearest neighbour type algorithms and rely on the choice of so-called summary statistics, distances between datasets and a tolerance threshold. Recently, methods combining ABC with more complex machine learning algorithms have been proposed to mitigate the impact of these “user-choices”. In this talk, I will present you the first, to our knowledge, ABC method completely free of summary statistics, distance, and tolerance threshold. Moreover, in contrast with usual generalisations of the ABC method, it associates a confidence interval (having a proper frequentist marginal coverage) with the posterior mean estimation (or other moment-type estimates). This method, named ABCD-Conformal, uses a neural network with Monte Carlo Dropout to provide an estimation of the posterior mean (or other moment type functionals), and conformal theory to obtain associated confidence sets. I will compare its performances with other ABC methods on several examples, and show you that it is efficient for estimating multidimensional parameters, while being “amortised”.
Keywords: simulation-based inference, approximate Bayesian computation, neural posterior estimation, convolutional neural networks, dropout, conformal prediction
statistical accuracy of neural posterior and likelihood estimation
Posted in pictures, Running, Statistics, Travel, University life with tags ABC, ABC consistency, Approximate Bayesian computation, Australia, Bayesian synthetic likelihood, Biometrika, Brisbane, Monash University, neural density estimator, neural posterior estimation, QUT, St Kilda, summary statistics on March 17, 2025 by xi'an
As I have been aiming at mentioning this news for quite a while, David Frazier, Ryan Kelly, Christopher Drovandi, and David Warne arXived last November a paper that parallels our paper (with David and Gael) on ABC consistency and some earlier papers of theirs for synthetic likelihood in the case of neural posterior approximations, under similar conditions (see, e.g., Assumptions 1 and 2), with potential reduced computational cost in some situations.
“NLE requires additional MCMC steps to produce a posterior approximation, whereas NPE produces a posterior approximation directly and does not require any additional sampling”
Convergence is achieved when the neural learning size grows fast enough with the sample size. And when the tolerance decreases fast enough with respect to the convergence rate of the summary statistic. Two options are possible, that is either approximating the likelihood and then exploiting this approximation in an MCMC algorithm, or directly approximating the posterior distribution, as a function of of the summary statistic Sn (rather than for the observed S⁰n), with arguments favouring the second option.
“if the intractable posterior Π(· | Sn) is asymptotically Gaussian a nd calibrated, then so long as νnγN = o(1), the NPE is also asymptotically Gaussian and calibrated”
where γN denotes the rate at which the neural approximation of the posterior converges to the ideal posterior (for the Kullback-Leibler divergence) in N the size of the learning sample. And νn is the rate of convergence of the statistic Sn to its asymptotic mean. The convergence result does not make explicit assumptions on the class of neural posteriors, but it requires that the observed statistic must fit within the range of the simulated values (a possibility illustrated in the paper with an MA(2) model that was already used in several of our papers (as I noticed when giving an ABC masterclass in Warwick this very week).
“While neural methods and normalizing flows are common choices for the approximating class Q, the diversity of such methods, along with their complicated tuning and training regimes, makes establishing theoretical results on the rate of convergence, γN, difficult”
Under stronger and hard to check assumptions, namely on the minimaxity of the posterior density estimator within the class of locally β-Hölder functions, they recover a closed form γN . Which unravels how N should be chosen (with a surprising addition of the dimensions of the parameter θ and of the summary Sn. With a resulting explosion in the theoretical minimal value of N one should use. (And decent performances of the method with smaller values of N!) Concerning minimaxity, I have no intuition how this impacts the sparseness (lack thereof) of the neural networks that can be used.
I am wondering at strategies to remove superfluous statistics since their dimension matters so much and in detecting or evaluating the misspecification (or its complement, the compatibility, as discussed on page 31). But all in all this paper represents a massive addition to the consistency results for approximate Bayesian inference methods!
