Archive for Bayesian model averaging

Bayesian workflow [book review]

Posted in Books, R, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on October 8, 2026 by xi'an


“This original, thought-provoking, and transforming, book is much much more than an implementation manual for Bayesian Data Analysis, even though it shares almost the same perspective. (The first sentence of the book states that the authors’ `conceptions of statistical practice, and of Bayesian statistics, have changed over the years’.) By providing a modus vivendi for undertaking Bayesian modelling from scratch in realistic settings where models are not magicked out of the blue, the authors explicit and rationalise the many steps required by such a bottom-up modelling protocol (`not a checklist, not a cookbook’, and not a flowchart!) in real situations. The contents read very well and very smoothly, with a seamless conjunction of intuition, modelling advices, computational details, and comparison tools. While unsurprisingly Bayesian, the perspective adopted therein remains both open and inclusive, with a welcome humility about the limitations and challenges of Bayesian workflows. This book should thus appeal to and profit a wide variety of readers, as providing guidance through an extensive collection of highly detailed examples, with shared code and exercises.”

This book proposes a modus vivendi for Bayesian modelling in applied, realistic Bayesian analysis, where models are not magicked out of the blue. It thus emphases iterative model building, model checking, computational troubleshooting, and simulated-data experimentation, filling a gap that looks glaring in retrospect. It particularly targets users and developers of Stan, with code excerpts in R and Stan. It consists of four parts:

  1. background on Bayesian methods and computational tools;
  2. the Bayesian workflow proper, namely building a statistical model from its components, together with its assessment tools;
  3. the computational aspects of fitting models, diagnosing convergence and assessing calibration;
  4. case studies.

I was eagerly waiting for the book, as I knew Andrew, Aki, and Richard had been working on it for a few years. (The quote above is the blurb I wrote upon request from the publisher.)

The tenets of BaWoFlo—if I may resort to this acronym!—are (i) fitting multiple models, (ii) applying methods repeatedly, and (iii) resorting to simulated-data experiments, which should not come as a surprise to readers of BDA. As noted in the introduction, the protocol exposed therein can also benefit non-Bayesian experimenters. This agrees with the highly moderate, “M-open”, agnostic approach to Bayesianism adopted by the authors (“there is no safe haven”). I also welcome and share their humble perspective about the limitations and challenges of Bayesian workflows.

Examples are treated in full detail, with successive modelling and computational choices profusely commented, which is a big plus for such a practical book. This starts as early as Chapter 4, with a multiple-choice exam example. Indeed, there cannot be general principles or a generic theory that would make the approach foolproof. See, e.g., “A data model is not just a ‘likelihood’” (p.70), as when the data model is not fully generative. I very much liked the section on choosing priors (5.6), and the very rich graphs (see, e.g., Chapter 8) for assessing the impact of prior and likelihood, as well as for predictive checks. In coherent continuation of the authors’ earlier work, the book advocates LOO methods and model stacking rather than model averaging. (With a surprisingly anti-Ockham perspective in Section 9.7.)

The MCMC coverage is unsurprising, with \(\hat R\) at the forefront. Chapter 12, on using fast experiments to detect fitting or computational issues, is very nice. The book builds on the immense corpus of work achieved by the authors over the decades (for the most senior ones!). By contrast, the chapter on approximate solutions (13) is way too short, and the same goes for those on calibration and software development.

The book is very US-centric, unsurprisingly given Andrew’s focus on political science. Some sections are reminiscent of Andrew’s blog entries (or the opposite). The (football) World Cup example was initiated when Andrew was in France, during the 2014 World Cup, and as a result (?) the names of the teams are in French! One chapter also reanalyses the birthdate data displayed on the cover of BDA.

Mileage varies on the applied chapters, depending on the example. A dog chapter is followed by a cat chapter! Not that the (stat)dog experiment was in any way enjoyable, especially for the dogs. Maybe the cats were running it! And then come chapters on roaches and sharks. There is also a frightening flowchart (Fig. 2.1)! And the book ends with an appendix on going through BDA to better understand BaWoFlo

[The usual disclaimer applies, namely that this review is likely to appear later in CHANCE, in my book reviews column.]

ISBA 2026²

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , on July 1, 2026 by xi'an

One morning session on optimal transport after first-hand witnessing the impressive ballet of orderly lines entering the subway at Nagoya Station (and a very early run along the river and a high humidity rate, hence the picture of empty street at 5am). With Hugo Lavenant exhibiting optimal rates for Bayesian nonparametrics using Wasserstein distances, Pierre Jacob coupling MCMC chains, and Anya Katsevich investigating non-Gaussian asymptotic distributions in high dimensions (beyond Bernstein-von Mises). Then I tried to attend the session on Bayesian Uncertainty Quantification and Posterior Sampling for Large-Scale Generative Models, but it proved too popular for the number of seats, and I ended up discussing with others. After a nap related to my jetlag induced, early, rise I went back to chair Sid Chib’s Foundation Lecture, where he discussed the use of an orbit of models in Bayesian model choice, rather than the (MAP) most likely one. Based on their 2018 JASA paper which I already discussed in Paris with Anna Simoni presenting. Hence reminding me of points I presumably already made, from the issue of having too many models to realistically explore to constructing coherent priors across them, with the fractional, empirical, proxy of using a (same) fraction of the sample as a learning sample, to more philosophical issues like missing a utility function about having to chose a model, especially with all models being wrong, missing an uncertainty quantification on the evidence itself, rather than using most likely models (MAP!), called an orbit by Sid (which requires some calibration). Since the uncertainty represented by the sample induces an uncertainty in the ranking of models. And a most appropriate, almost local, occurrence of the Rashomon principle!!! And I finished the day mixing with many friends in the poster session¹, where Darren Wraith presented our ongoing work on novel, adaptive, importance, sampling.

JSM 2024, Portland, Day 3

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on August 9, 2024 by xi'an

Bayesian contributed session as the first round of the third day (with a choice of five parallel sessions featuring Bayesian topics!!, actually easier to pick than among the following eight parallel sessions of the 10:30 schedule!!!), with a talk by Tahir Ekin on adversarial outlier detection that could connect with our Oceaner(c) privacy concerns. Then one involving spike & slab (a theme to figure prominently in this special day!!) in mixed response models by Sameer Deshpande, seeking a (unBayesian!) MAP for a latent variable model by Monte Carlo EM. Followed by a talk by Yunyi Shen on completely random measures for estimating the (distribution of the) number of species in heterogeneous populations. Next, Valentin Zulj on (frequentist rather than) Bayesian stacking, on estimating optimal weights for model averaging (which should be posterior probabilities in a pure Bayesian mindframe), including a score function that could lead to generalised Bayesian inference on said weights. Finishing with a talk by Chaegeun Song on correcting Bayesian credible sets towards (frequentist, again!!!) exact coverage for classification (which reminded me of my very first paper with George on correcting frequentist confidence for Binomial observations). With which I could not really engage as seeking a specific coverage level did not seem relevant, imho, but I appreciated the wheel plot representation.My second morn session was about modern (what else?!) sampling algorithms, although I spent the first dozen minutes wondering whether or not I had entered the wrong room. Until Tianhao Wang focussed on Thompson sampling for bandits. It did prove far enough from my interest for my (sleep deprived) attention to drift too quickly. Only the talk by Yuchen Wu on a spike & slab (as suits the day!) challenge captured enough this wandering attention. Crossing further into my realm of primary topics by considering a target distribution that is a product of distributions. But I did not get from her presentation how a product measure decomposition was inducing higher efficiency (and did not find answers within the arXived preprint). Unless it exploited specific features of the target, like conditional independence between the components. The last talk was by Brice Huang on sampling low temperature Gibbs measures using stochastic localisation.

After coming upon a row of food trucks across the conference centre and being unfairly attracted by an Ethiopian injera picture into a terrible wrap, I returned for the Skeptical about AI session, just a few minutes late, only to find accessing the session was impossible! Quite sad to miss the presentations and the arguments (even though I had heard a previous talk by Genevera Allen when visiting Rutgers two years ago). As a second best, I then joined the recent (of course!) Advances in Bayesian Computation (aka ABC?!) session with a medley of topics, including a data subset versus data sketching model reduction by Sudipto Saha. Which could have consequences on our privacy strategies. And marginal evidence estimation for the Bayesian Lasso by Christopher Hans while avoiding data completion. And another latent variable model with a sequential variational Bayes approach by Bao Anh Vu, using at one point Cappé et al. (2005) EM-based approximation to the log likelihood gradient. Finishing by a back-to-the-future talk by Luke Duttweiler on MCMC convergence diagnostics. Comparing several chains via proximity maps that themselves require some preliminary knowledge about the MCMC kernel. (Nice title though, “the traceplot thickens”!)The crux of the day was however the 2024 COPSS Award ceremony with several friends featuring among the recipients, Danielle Durante for the Emerging Leaders Award, Regina Liu for the Elizabeth L. Scott Award and Veronika Rockova for the Presidents’ Award. Congrats!!!



simulating the pandemic

Posted in Books, Statistics with tags , , , , , , , , , , , on November 28, 2020 by xi'an

Nature of 13 November has a general public article on simulating the COVID pandemic as benefiting from the experience gained by climate-modelling methodology.

“…researchers didn’t appreciate how sensitive CovidSim was to small changes in its inputs, their results overestimated the extent to which a lockdown was likely to reduce deaths…”

The argument is essentially Bayesian, namely rather than using a best guess of the parameters of the model, esp. given the state of the available data (and the worse for March). When I read

“…epidemiologists should stress-test their simulations by running ‘ensemble’ models, in which thousands of versions of the model are run with a range of assumptions and inputs, to provide a spread of scenarios with different probabilities…”

it sounds completely Bayesian. Even though there is no discussion of the prior modelling or of the degree of wrongness of the epidemic model itself. The researchers at UCL who conducted the multiple simulations and the assessment of sensitivity to the 940 various parameters found that 19 of them had a strong impact, mostly

“…the length of the latent period during which an infected person has no symptoms and can’t pass the virus on; the effectiveness of social distancing; and how long after getting infected a person goes into isolation…”

but this outcome is predictable (and interesting). Mentions of Bayesian methods appear at the end of the paper:

“…the uncertainty in CovidSim inputs [uses] Bayesian statistical tools — already common in some epidemiological models of illnesses such as the livestock disease foot-and-mouth.”

and

“Bayesian tools are an improvement, says Tim Palmer, a climate physicist at the University of Oxford, who pioneered the use of ensemble modelling in weather forecasting.”

along with ensemble modelling, which sounds a synonym for Bayesian model averaging… (The April issue on the topic had also Bayesian aspects that were explicitely mentionned.)

Markov melding

Posted in Books, Statistics, University life with tags , , , on July 2, 2020 by xi'an

“An alternative approach is to model smaller, simpler aspects of the data, such that designing these submodels is easier, then combine the submodels.”

An interesting paper by Andrew Manderson and Robert Goudie I read on arXiv on merging (or melding) several models together. With different data and different parameters. The assumption is one of a common parameter φ shared by all (sub)models. Since the product of the joint distributions across the m submodels involves m replicates of φ, the melded distribution is the product of the conditional distributions given φ, times a common (or pooled) prior on φ. Which leads to a perfectly well-defined joint distribution provided the support of this pooled prior is compatible with all conditionals.

The MCMC aspects of such a target are interesting in that the submodels can easily be exploited to return proposal distributions on their own parameters (plus φ). Although the notion is fraught with danger when considering a flat prior on φ, since the posterior is not necessarily well-defined. Or at the very least unrelated with the actual marginal posterior. This first stage is used to build a particle approximation to the posterior distribution of φ, exploited in the later simulation of the other subsample parameters and updates of φ. Due to the rare availability of the (submodel) marginal prior on φ, it is replaced in the paper by a kernel density estimate. Not a great idea as (a) it is unstable and (b) the joint density is costly, while existing! Which brings the authors to set a goal of estimating a ratio. Of the same marginal density in two different values of φ. (Not our frequent problem of the ratio of different marginals!) They achieve this by targeting another joint, using a weight function both for the simulation and the kernel density estimation… Requiring the calibration of the weight function and the production of a biased estimate of the ratio.

While the paper concentrates very much on computational improvements, including the possible recourse to unbiased MCMC, I also feel it is missing on the Bayesian aspects, since the construction of the multi-level Bayesian model faces many challenges. In a sense this is an alternative to our better together paper, where cuts are used to avoid the duplication of common parameters.