Archive for the R Category

Bayesian workflow [book review]

Posted in Books, R, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on October 8, 2026 by xi'an


“This original, thought-provoking, and transforming, book is much much more than an implementation manual for Bayesian Data Analysis, even though it shares almost the same perspective. (The first sentence of the book states that the authors’ `conceptions of statistical practice, and of Bayesian statistics, have changed over the years’.) By providing a modus vivendi for undertaking Bayesian modelling from scratch in realistic settings where models are not magicked out of the blue, the authors explicit and rationalise the many steps required by such a bottom-up modelling protocol (`not a checklist, not a cookbook’, and not a flowchart!) in real situations. The contents read very well and very smoothly, with a seamless conjunction of intuition, modelling advices, computational details, and comparison tools. While unsurprisingly Bayesian, the perspective adopted therein remains both open and inclusive, with a welcome humility about the limitations and challenges of Bayesian workflows. This book should thus appeal to and profit a wide variety of readers, as providing guidance through an extensive collection of highly detailed examples, with shared code and exercises.”

This book proposes a modus vivendi for Bayesian modelling in applied, realistic Bayesian analysis, where models are not magicked out of the blue. It thus emphases iterative model building, model checking, computational troubleshooting, and simulated-data experimentation, filling a gap that looks glaring in retrospect. It particularly targets users and developers of Stan, with code excerpts in R and Stan. It consists of four parts:

  1. background on Bayesian methods and computational tools;
  2. the Bayesian workflow proper, namely building a statistical model from its components, together with its assessment tools;
  3. the computational aspects of fitting models, diagnosing convergence and assessing calibration;
  4. case studies.

I was eagerly waiting for the book, as I knew Andrew, Aki, and Richard had been working on it for a few years. (The quote above is the blurb I wrote upon request from the publisher.)

The tenets of BaWoFlo—if I may resort to this acronym!—are (i) fitting multiple models, (ii) applying methods repeatedly, and (iii) resorting to simulated-data experiments, which should not come as a surprise to readers of BDA. As noted in the introduction, the protocol exposed therein can also benefit non-Bayesian experimenters. This agrees with the highly moderate, “M-open”, agnostic approach to Bayesianism adopted by the authors (“there is no safe haven”). I also welcome and share their humble perspective about the limitations and challenges of Bayesian workflows.

Examples are treated in full detail, with successive modelling and computational choices profusely commented, which is a big plus for such a practical book. This starts as early as Chapter 4, with a multiple-choice exam example. Indeed, there cannot be general principles or a generic theory that would make the approach foolproof. See, e.g., “A data model is not just a ‘likelihood’” (p.70), as when the data model is not fully generative. I very much liked the section on choosing priors (5.6), and the very rich graphs (see, e.g., Chapter 8) for assessing the impact of prior and likelihood, as well as for predictive checks. In coherent continuation of the authors’ earlier work, the book advocates LOO methods and model stacking rather than model averaging. (With a surprisingly anti-Ockham perspective in Section 9.7.)

The MCMC coverage is unsurprising, with \(\hat R\) at the forefront. Chapter 12, on using fast experiments to detect fitting or computational issues, is very nice. The book builds on the immense corpus of work achieved by the authors over the decades (for the most senior ones!). By contrast, the chapter on approximate solutions (13) is way too short, and the same goes for those on calibration and software development.

The book is very US-centric, unsurprisingly given Andrew’s focus on political science. Some sections are reminiscent of Andrew’s blog entries (or the opposite). The (football) World Cup example was initiated when Andrew was in France, during the 2014 World Cup, and as a result (?) the names of the teams are in French! One chapter also reanalyses the birthdate data displayed on the cover of BDA.

Mileage varies on the applied chapters, depending on the example. A dog chapter is followed by a cat chapter! Not that the (stat)dog experiment was in any way enjoyable, especially for the dogs. Maybe the cats were running it! And then come chapters on roaches and sharks. There is also a frightening flowchart (Fig. 2.1)! And the book ends with an appendix on going through BDA to better understand BaWoFlo

[The usual disclaimer applies, namely that this review is likely to appear later in CHANCE, in my book reviews column.]

Round-robin (with Claude)

Posted in Books, Kids, R, Statistics, Wines with tags , , , , , , , , , , , , on August 15, 2026 by xi'an

A few days ago I had a coffee in Paris with my long-time friend (and former Statistics & Computing editor) Gilles Celeux, and he mentioned me stopping solving and posting maths puzzles like those weekly published by Le Monde. They have indeed vanished with the retirement of the authors, but Gilles added that the arrival of LLMs would have made the exercise moot. I disagreed as (i) the fun of solving the puzzle on my own  has not gone away and (ii) the pedagogical appeal of the puzzle and its resolution remains. As the next Fiddler puzzle arrived in my mailbox, my resolution was put to the test (contrariwise to the previous entry, which did not require massive computations):

The Fiddler League consists of two teams. Over a season, they play each other 162 times. Each team has an equal chance of winning each game, and the results of games are independent. Over the season, on average, how many games would you expect the team with the better record to have won?

I started on the wrong foot with E[X|X≥81] when X is Bin(162,½), equal to 85.77 (either directly or with a Normal approximation), which differs from my second thought, E[max(X,162-X)]=86.07 (either directly or with a Normal approximation), which is larger because of the reflection produced by max. While the first computation was manageable, the second one seemed to involve simulation and I caved in prompting Claude, which provided the answer along with the connection

E[max(X,162−X)]=E[X∣X≥81](1+p81​)−81p81​

After some expansion, the League boasts 30 teams. Over a season, each team plays each other team five times. (Each team plays a total of 145 games.) Again, each team has an equal chance of winning each game, and the results of games are independent. Over the season, on average, how many games would you expect the team with the best record to have won?

The best record is max(Xi) with each of the 30 Xi‘s a sum of 29 Yij and the Yij=5-Yji distributed as Bin(5,½). The Xi‘s are thus Bin (145,½) but dependent. While I could not figure out a closed form answer for the expectation, a direct Monte Carlo resolution is obviously feasible, with Claude (rather than me) running it over 400 million repetitions, but a 30 dimensional Normal approximation exploiting the correlation of 1/29 between the components leads to roughly 85 as the expected value. (Again computed by a Claudicant simulation.)

While the conclusion that the Normal approximation is pretty accurate with so many terms in the Binomial variates is quite unsurprising, Claude saves me coding time without ruining the puzzle altogether. (And Gemini made me aware that the name of the café where Gilles and I regularly meet, L’Écir, is an Auvergne noun for a local, dangerous, mountain blizzard! Thus linking the place to the foundation of the café by Auvergne expatriates…)

R’ousseeuw²⁶ prize!

Posted in R, Statistics, University life with tags , , , , , , , , , , on June 29, 2026 by xi'an

Great news that a major Statistics prize like the Rousseeuw Prize goes to the R Core Team, esp. Brian Ripley (University of Oxford), Martin Mächler (ETH Zürich), Kurt Hornik (WU Wien), Peter Dalgaard (Copenhagen Business School), and Luke Tierney (University of Iowa). R is indeed a unique phenomenon, where open-source and open-access has been developed by and for the statistics community. Which is about to release R version 4.6.1 (Happy Hop) version.

Thanks to the R Core Team (and congrats!). Half of the Prize goes to the other members of the Team.

“The international and independent jury, appointed by the King Baudouin Foundation, has recognised the groundbreaking work of five members of the R Core Team who have been awarded the Rousseeuw Prize for Statistics. The international award, which recognises major contributions to statistical research, honours their nearly three decades of unpaid work building R, the open-source language that has become the common foundation of modern statistical computing.

Statistics is everywhere. It determines whether a new medicine is safe enough to reach patients, monitors risk in financial markets, and tracks how diseases spread. R is the tool that made it accessible to everyone. The language that the R Core Team built is trusted by institutions including the US Food and Drug Administration, the European Central Bank, the Bank of England, and major pharmaceutical companies worldwide.”

Bayesian Workflow [cover]

Posted in Books, pictures, R, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , on April 25, 2026 by xi'an

Ah great, the new book on Bayesian workflow by Andrew Gelman, Aki Vehtari, Richard McElreath I knew they were working on is about to appear!  With entries from several coauthors and half of the chapters on case studies. I have not (yet) looked at its contents in detail…

“Approximating evidence via bounded harmonic means” is out! [in Statistics and Computing]

Posted in Books, R, Statistics, University life with tags , , , , , , , , , , , on April 24, 2026 by xi'an