A few days before the January OWABI, I read through Simon Kucharsky’s and Paul Bürkner’s paper, arXived on 17 January. Which proposed an amortized Bayesian inference (ABI) method, even though the ABI is not the same as in OWABI! The motivation for their work is to start from a (standard) mixture model where the components are not analytically tractable (but still parameterised). But a generative model nonetheless. As in the earlier reviewed paper (which was arXived on the same day), by MEJ Newman, the dual representation of the joint posterior p(θ,z|x) as p(z|x,θ)p(θ|x) and p(θ|z,x)p(z|x) is (over?) emphasized (albeit unclearly why!). ABI uses neural networks and more specifically normalising flows to approximate the posterior p(θ|x) from prior predictive samples (θ,x) (as in ABC), and then directly exploit the invertibility of said flows to generate from this approximate posterior. One interesting aspect of the modelling is the derivation of summary statistics in the design of the network, albeit mixture posteriors do not allow for dimension-reduced (Bayes) sufficient statistics (and a contradictory sentence that conditioning on the summaries “does not alter the target posterior”, p7). The resulting approximate posterior generator proves much much faster than running an MCMC, obviously, and furthermore adapt to handling a sequence of datasets. A second network is constructed to approximate p(z|x,θ), using the same summaries. The network parameters are estimated through losses, rather than in a Bayesian manner, with a default Kullback-Leibler version (18). I also fail to understand why the networks are trained over unconstrained parameters when all parameters could become unconstrained when using the adequate parameterisation. And am fairly surprised at the regression towards the ill-fated step of using ordered parameters to avoid label switching… But the main quandary remains the issue of assessing the approximation effect, despite experiments aiming at pacifying such worries. And similarities with Stan and BayesFlow.
Archive for sufficiency
amortized Bayesian mixture model
Posted in Books, Statistics, University life with tags ABC, amortization, amortized Bayesian inference, Approximate Bayesian computation, approximate Bayesian inference, dimension reduction, doubly intractable posterior, label switching, mixtures of distributions, normalising flow, OWABI, reparameterisation, STAN, sufficiency, summary statistics, webinar on February 7, 2025 by xi'anInsufficient Gibbs sampling [at COMPSTAT 2024]
Posted in pictures, Statistics, Travel, University life with tags AI, Bayes(Pharma), bridge sampling, Compstat 2024, conference, Germany, Gießen, Hesse, insufficient statistic, lecture, pharmacokinetics, pharmacometrics, sufficiency, Université Paris Dauphine on August 29, 2024 by xi'an
The COMPSTAT 2024 programme proved somewhat remote from my interests, far from the 1998 version I attended that was buzzing and brimming with MCMC sessions! (Correlatively, apart from the speakers in my own session, I hardly knew anyone there.) I however [missed a Bayesian session as I] attended a session on change-point detection, with a talk by Ziyang Yang (Lancaster University) on an acceleration proposal when the data size is too much, by aggregating similar datapoints, for which I would like to see a theoretical analysis of the impact of this aggregation on the quality (and consistency) of the attached Bayes factor. In my own session, I did not feel overly comfortable with the presentation of a commercial [i.e., for sale] software for pharmacokinetics.
I also enjoyed spending the day in Gießen, from running to the top of nearby Burg Glieberg at sunrise to swimming outdoor in the local 50m Freibad.

learning optimal summary statistics
Posted in Books, pictures, Statistics with tags ABC, Approximate Bayesian computation, Bayesian inference, Fisher information, Kullback-Leibler divergence, neural density estimator, Normandy, quiz, sufficiency, summary statistics on July 27, 2022 by xi'an“Despite the pursuit of the holy grail of sufficient statistics, most applications will have to settle for the weakest concept of optimal statistics.”
Quiz #1: How does Bayes sufficiency [which preserves the posterior density] differ from sufficiency [which preserves the likelihood function]?
Quiz #2: How does Fisher-information sufficiency [which preserves the information matrix] differ from standard sufficiency [which preserves the likelihood function]?
Read a recent arXival by Till Hoffmann and Jukka-Pekka Onnela that I frankly found most puzzling… Maybe due to the Norman train where I was traveling being particularly noisy.
The argument in the paper is to find a summary statistic that minimises the [empirical] expected posterior entropy, which equivalently means minimising the expected Kullback-Leibler distance to the full posterior. And maximizing the mutual information between parameters θ and summaries t(.). And maximizing the expected surprise. Which obviously requires breaking the sample into iid components and hence considering the gain brought by a specific transform of a single observation. The paper also contains a long comparison with other criteria for choosing summaries.
“Minimizing the posterior entropy would discard the sufficient statistic t such that the posterior is equal to the prior–we have not learned anything from the data.”
Furthermore, the expected aspect of the criterion takes us away from a proper Bayes analysis (and exhibits artifacts as the one above), which somehow makes me question the relevance of comparing entropies under different distributions. It took me a long while to realise that the collection of summaries was set by the user and quite limited. Like a neural network representation of the posterior mean. And the intractable posterior is further approximated by a closed-form function of the parameter θ and of the summary t(.). Using there a neural density estimator. Or a mixture density network.
day one at ISBA 22
Posted in pictures, Statistics, Travel, University life with tags ABC, ABC-PMC, arXiv, BFF Statistics Conference, bistronomie, Bruno de Finetti, Canada, copulas, distilled importance sampling, evidence, exchangeability, fiducial inference, John Maynard Keynes, Korean food, Montréal, normalising flow, optimal transport, population demography, Québec, R.A. Fisher, Rademacher complexity, restaurant, Roe v. Wade, SMC-ABC, sufficiency, Supreme Court, United Nations Population Fund on June 29, 2022 by xi'an
Started the day with a much appreciated swimming practice in the [alas warm⁺⁺⁺] outdoor 50m pool on the Island with no one but me in the slooow lane. And had my first ride with the biXi system, surprised at having to queue behind other bikes at red lights! More significantly, it was a great feeling to reunite at last with so many friends I had not met for more than two years!!!
My friend Adrian Raftery gave the very first plenary lecture on his work on the Bayesian approach to long-term population projections, which was recently a work censored by some US States, then counter-censored by the Supreme Court [too busy to kill Roe v. Wade!]. Great to see the use of Bayesian methods validated by the UN Population Division [with at least one branch of the UN
Stephen Lauritzen returning to de Finetti notion of a model as something not real or true at all, back to exchangeability. Making me wonder when exchangeability is more than a convenient assumption leading to the Hewitt-Savage theorem. And sufficiency. I mean, without falling into a Keynesian fallacy, each point of the sample has unique specificities that cannot be taken into account in an exchangeable model. Nice to hear some measure theory, though!!! Plus a comment on the median never being sufficient, recouping an older (and presumably not original) point of mine. Stephen’s (or Fisher’s?) argument being that the median cannot be recursively computed!
Antonietta Mira and I had our ABC session this afternoon with Cecilia Viscardi, Sirio Legramanti, and Massimiliano Tamborino (Warwick) as speakers. Cecilia linked ABC with normalising flows, in collaboration with Dennis Prangle (whose earlier paper on this connection was presented as the first One World ABC seminar). Thus using past simulations to approximate the posterior by a neural network, possibly with a significant increase in computing time when compared with more rudimentary SMC-ABC methods in larger dimensions. Sirio considered summary-free ABC based on discrepancies like Rademacher complexity. Which more or less contains MMD, Kullback-Leibler, Wasserstein and more, although it seems to be dependent on the parameterisation of the observations. An interesting opening at the end was that this approach could apply to non iid settings. Massi presented a paper coauthored with Umberto that had just been arXived. On sequential ABC with a dependence on the summary statistic (hence guided). Further bringing copulas into the game, although this forces another choice [for the marginals] in the method.
Tamara Broderick talked about a puzzling leverage effect of some observations in economic studies where a tiny portion of individuals may modify the significance or the sign of a coefficient, for which I cannot tell whether the data or the reliance on statistical significance are to blame. Robert Kohn presented mixture-of-Gaussian copulas [not to be confused with mixture of Gaussian-copulas!] and Nancy Reid concluded my first [and somewhat exhausting!] day at ISBA with a BFF talk on the different statistical paradigms take on confidence (for which the notion of calibration seems to remain frequentist).
Side comments: First, most people in the conference are wearing masks, which is great! Also, I find it hard to read slides from the screen, which I presume is an age issue (?!) Even more aside, I had Korean lunch in a place that refused to serve me a glass of water, which I find amazing.
set-valued sufficient statistic
Posted in Books, Kids, Statistics with tags best unbiased estimator, cross validated, sufficiency, sufficient statistics, theory of statistics, UMVUE, uniform distribution on June 18, 2022 by xi'anWhile the classical definition of a statistic is one of a real valued random variable or vector, less usual situations call for broader definitions… For instance, in an homework problem from Mark Schervish’s Theory of Statistics, a sample from the uniform distribution of a ball of unknown centre θ and radius ς is associated with the convex hull of said sample as “sufficient statistic”, albeit the object being a set. Similarly, if the radius ς is known, the set made of the intersection of all the balls of radius ς centred at the observations is sufficient, in that the likelihood is constant for θ inside and zero outside. As discussed in this X validated question, this does not define an optimal estimator of the center θ, while Pitman’s best location equivariant does, while the centre of this sufficient set, but it is not sufficient as a statistic and is not necessarily the MVUE, if unbiased.
