Archive for model checking

Bayesian workflow [book review]

Posted in Books, R, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on October 8, 2026 by xi'an


“This original, thought-provoking, and transforming, book is much much more than an implementation manual for Bayesian Data Analysis, even though it shares almost the same perspective. (The first sentence of the book states that the authors’ `conceptions of statistical practice, and of Bayesian statistics, have changed over the years’.) By providing a modus vivendi for undertaking Bayesian modelling from scratch in realistic settings where models are not magicked out of the blue, the authors explicit and rationalise the many steps required by such a bottom-up modelling protocol (`not a checklist, not a cookbook’, and not a flowchart!) in real situations. The contents read very well and very smoothly, with a seamless conjunction of intuition, modelling advices, computational details, and comparison tools. While unsurprisingly Bayesian, the perspective adopted therein remains both open and inclusive, with a welcome humility about the limitations and challenges of Bayesian workflows. This book should thus appeal to and profit a wide variety of readers, as providing guidance through an extensive collection of highly detailed examples, with shared code and exercises.”

This book proposes a modus vivendi for Bayesian modelling in applied, realistic Bayesian analysis, where models are not magicked out of the blue. It thus emphases iterative model building, model checking, computational troubleshooting, and simulated-data experimentation, filling a gap that looks glaring in retrospect. It particularly targets users and developers of Stan, with code excerpts in R and Stan. It consists of four parts:

  1. background on Bayesian methods and computational tools;
  2. the Bayesian workflow proper, namely building a statistical model from its components, together with its assessment tools;
  3. the computational aspects of fitting models, diagnosing convergence and assessing calibration;
  4. case studies.

I was eagerly waiting for the book, as I knew Andrew, Aki, and Richard had been working on it for a few years. (The quote above is the blurb I wrote upon request from the publisher.)

The tenets of BaWoFlo—if I may resort to this acronym!—are (i) fitting multiple models, (ii) applying methods repeatedly, and (iii) resorting to simulated-data experiments, which should not come as a surprise to readers of BDA. As noted in the introduction, the protocol exposed therein can also benefit non-Bayesian experimenters. This agrees with the highly moderate, “M-open”, agnostic approach to Bayesianism adopted by the authors (“there is no safe haven”). I also welcome and share their humble perspective about the limitations and challenges of Bayesian workflows.

Examples are treated in full detail, with successive modelling and computational choices profusely commented, which is a big plus for such a practical book. This starts as early as Chapter 4, with a multiple-choice exam example. Indeed, there cannot be general principles or a generic theory that would make the approach foolproof. See, e.g., “A data model is not just a ‘likelihood’” (p.70), as when the data model is not fully generative. I very much liked the section on choosing priors (5.6), and the very rich graphs (see, e.g., Chapter 8) for assessing the impact of prior and likelihood, as well as for predictive checks. In coherent continuation of the authors’ earlier work, the book advocates LOO methods and model stacking rather than model averaging. (With a surprisingly anti-Ockham perspective in Section 9.7.)

The MCMC coverage is unsurprising, with \(\hat R\) at the forefront. Chapter 12, on using fast experiments to detect fitting or computational issues, is very nice. The book builds on the immense corpus of work achieved by the authors over the decades (for the most senior ones!). By contrast, the chapter on approximate solutions (13) is way too short, and the same goes for those on calibration and software development.

The book is very US-centric, unsurprisingly given Andrew’s focus on political science. Some sections are reminiscent of Andrew’s blog entries (or the opposite). The (football) World Cup example was initiated when Andrew was in France, during the 2014 World Cup, and as a result (?) the names of the teams are in French! One chapter also reanalyses the birthdate data displayed on the cover of BDA.

Mileage varies on the applied chapters, depending on the example. A dog chapter is followed by a cat chapter! Not that the (stat)dog experiment was in any way enjoyable, especially for the dogs. Maybe the cats were running it! And then come chapters on roaches and sharks. There is also a frightening flowchart (Fig. 2.1)! And the book ends with an appendix on going through BDA to better understand BaWoFlo

[The usual disclaimer applies, namely that this review is likely to appear later in CHANCE, in my book reviews column.]

Bayesian Workflow [cover]

Posted in Books, pictures, R, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , on April 25, 2026 by xi'an

Ah great, the new book on Bayesian workflow by Andrew Gelman, Aki Vehtari, Richard McElreath I knew they were working on is about to appear!  With entries from several coauthors and half of the chapters on case studies. I have not (yet) looked at its contents in detail…

Measuring statistical evidence using relative belief [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , on July 22, 2015 by xi'an

“It is necessary to be vigilant to ensure that attempts to be mathematically general do not lead us to introduce absurdities into discussions of inference.” (p.8)

This new book by Michael Evans (Toronto) summarises his views on statistical evidence (expanded in a large number of papers), which are a quite unique mix of Bayesian  principles and less-Bayesian methodologies. I am quite glad I could receive a version of the book before it was published by CRC Press, thanks to Rob Carver (and Keith O’Rourke for warning me about it). [Warning: this is a rather long review and post, so readers may chose to opt out now!]

“The Bayes factor does not behave appropriately as a measure of belief, but it does behave appropriately as a measure of evidence.” (p.87)

Continue reading →

posterior predictive p-values

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , on February 4, 2014 by xi'an

Bayesian Data Analysis advocates in Chapter 6 using posterior predictive checks as a way of evaluating the fit of a potential model to the observed data. There is a no-nonsense feeling to it:

“If the model fits, then replicated data generated under the model should look similar to observed data. To put it another way, the observed data should look plausible under the posterior predictive distribution.”

And it aims at providing an answer to the frustrating (frustrating to me, at least) issue of Bayesian goodness-of-fit tests. There are however issues with the implementation, from deciding on which aspect of the data or of the model is to be examined, to the “use of the data twice” sin. Obviously, this is an exploratory tool with little decisional backup and it should be understood as a qualitative rather than quantitative assessment. As mentioned in my tutorial on Sunday (I wrote this post in Duke during O’Bayes 2013), it reminded me of Ratmann et al.’s ABCμ in that they both give reference distributions against which to calibrate the observed data. Most likely with a multidimensional representation. And the “use of the data twice” can be argued for or against, once a data-dependent loss function is built.

“One might worry about interpreting the significance levels of multiple tests or of tests chosen by inspection of the data (…) We do not make [a multiple test] adjustment, because we use predictive checks to see how particular aspects of the data would be expected to appear in replications. If we examine several test variables, we would not be surprised for some of them not to be fitted by the model-but if we are planning to apply the model, we might be interested in those aspects of the data that do not appear typical.”

The natural objection that having a multivariate measure of discrepancy runs into multiple testing is answered within the book with the reply that the idea is not to run formal tests. I still wonder how one should behave when faced with a vector of posterior predictive p-values (ppp).

pospredThe above picture is based on a normal mean/normal prior experiment I ran where the ratio prior-to-sampling variance increases from 100 to 10⁴. The ppp is based on the Bayes factor against a zero mean as a discrepancy. It thus grows away from zero very quickly and then levels up around 0.5, reaching only values close to 1 for very large values of x (i.e. never in practice). I find the graph interesting because if instead of the Bayes factor I use the marginal (numerator of the Bayes factor) then the picture is the exact opposite. Which, I presume, does not make a difference for Bayesian Data Analysis, since both extremes are considered as equally toxic… Still, still, still, we are is the same quandary as when using any kind of p-value: what is extreme? what is significant? Do we have again to select the dreaded 0.05?! To see how things are going, I then simulated the behaviour of the ppp under the “true” model for the pair (θ,x). And ended up with the histograms below:

truepospredwhich shows that under the true model the ppp does concentrate around .5 (surprisingly the range of ppp’s hardly exceeds .5 and I have no explanation for this). While the corresponding ppp does not necessarily pick any wrong model, discrepancies may be spotted by getting away from 0.5…

“The p-value is to the u-value as the posterior interval is to the confidence interval. Just as posterior intervals are not, in general, classical confidence intervals, Bayesian p-values are not generally u-values.”

Now, Bayesian Data Analysis also has this warning about ppp’s being not uniform under the true model (u-values), which is just as well considering the above example, but I cannot help wondering if the authors had intended a sort of subliminal message that they were not that far from uniform. And this brings back to the forefront the difficult interpretation of the numerical value of a ppp. That is, of its calibration. For evaluation of the fit of a model. Or for decision-making…