Archive for CHANCE

Bayesian workflow [book review]

Posted in Books, R, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on October 8, 2026 by xi'an


“This original, thought-provoking, and transforming, book is much much more than an implementation manual for Bayesian Data Analysis, even though it shares almost the same perspective. (The first sentence of the book states that the authors’ `conceptions of statistical practice, and of Bayesian statistics, have changed over the years’.) By providing a modus vivendi for undertaking Bayesian modelling from scratch in realistic settings where models are not magicked out of the blue, the authors explicit and rationalise the many steps required by such a bottom-up modelling protocol (`not a checklist, not a cookbook’, and not a flowchart!) in real situations. The contents read very well and very smoothly, with a seamless conjunction of intuition, modelling advices, computational details, and comparison tools. While unsurprisingly Bayesian, the perspective adopted therein remains both open and inclusive, with a welcome humility about the limitations and challenges of Bayesian workflows. This book should thus appeal to and profit a wide variety of readers, as providing guidance through an extensive collection of highly detailed examples, with shared code and exercises.”

This book proposes a modus vivendi for Bayesian modelling in applied, realistic Bayesian analysis, where models are not magicked out of the blue. It thus emphases iterative model building, model checking, computational troubleshooting, and simulated-data experimentation, filling a gap that looks glaring in retrospect. It particularly targets users and developers of Stan, with code excerpts in R and Stan. It consists of four parts:

  1. background on Bayesian methods and computational tools;
  2. the Bayesian workflow proper, namely building a statistical model from its components, together with its assessment tools;
  3. the computational aspects of fitting models, diagnosing convergence and assessing calibration;
  4. case studies.

I was eagerly waiting for the book, as I knew Andrew, Aki, and Richard had been working on it for a few years. (The quote above is the blurb I wrote upon request from the publisher.)

The tenets of BaWoFlo—if I may resort to this acronym!—are (i) fitting multiple models, (ii) applying methods repeatedly, and (iii) resorting to simulated-data experiments, which should not come as a surprise to readers of BDA. As noted in the introduction, the protocol exposed therein can also benefit non-Bayesian experimenters. This agrees with the highly moderate, “M-open”, agnostic approach to Bayesianism adopted by the authors (“there is no safe haven”). I also welcome and share their humble perspective about the limitations and challenges of Bayesian workflows.

Examples are treated in full detail, with successive modelling and computational choices profusely commented, which is a big plus for such a practical book. This starts as early as Chapter 4, with a multiple-choice exam example. Indeed, there cannot be general principles or a generic theory that would make the approach foolproof. See, e.g., “A data model is not just a ‘likelihood’” (p.70), as when the data model is not fully generative. I very much liked the section on choosing priors (5.6), and the very rich graphs (see, e.g., Chapter 8) for assessing the impact of prior and likelihood, as well as for predictive checks. In coherent continuation of the authors’ earlier work, the book advocates LOO methods and model stacking rather than model averaging. (With a surprisingly anti-Ockham perspective in Section 9.7.)

The MCMC coverage is unsurprising, with \(\hat R\) at the forefront. Chapter 12, on using fast experiments to detect fitting or computational issues, is very nice. The book builds on the immense corpus of work achieved by the authors over the decades (for the most senior ones!). By contrast, the chapter on approximate solutions (13) is way too short, and the same goes for those on calibration and software development.

The book is very US-centric, unsurprisingly given Andrew’s focus on political science. Some sections are reminiscent of Andrew’s blog entries (or the opposite). The (football) World Cup example was initiated when Andrew was in France, during the 2014 World Cup, and as a result (?) the names of the teams are in French! One chapter also reanalyses the birthdate data displayed on the cover of BDA.

Mileage varies on the applied chapters, depending on the example. A dog chapter is followed by a cat chapter! Not that the (stat)dog experiment was in any way enjoyable, especially for the dogs. Maybe the cats were running it! And then come chapters on roaches and sharks. There is also a frightening flowchart (Fig. 2.1)! And the book ends with an appendix on going through BDA to better understand BaWoFlo

[The usual disclaimer applies, namely that this review is likely to appear later in CHANCE, in my book reviews column.]

renewable energy [book review]

Posted in Books, Kids, Travel with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on December 28, 2025 by xi'an

Renewable Energy (2nd edition, 2025), by Nick Jelley is part of the terrific “Very Short Introduction” series, which provides an expert introduction to a topic within 150 pages. (The OUP equivalent of the French series “Que sais-je?” that started in 1941.) The reason for the second edition, as provided by the author, is the “dramatic expansion, and fall in cost” of clean energy products. Unfortunately, it comes out just short of realising the magnitude of the backlash again renewable energy and fighting climate change launched by the second Trump administration and the ensuing added pressure on other countries to reduce further their fuel consumption.

The main sources of renewables are identified as wind, sun, and water, in a first chapter that operates as an historical recap on the evolution of energy sources and consumption. The second chapter stresses the need for renewable to fight global warming and to reverse climate change, with a rather vague discussion of the costs of producing energy from renewable. Chapter 3 focusses on (debatable) biomass, solar heat, and hydropower. Chapter 4 on wind power, with a few paragraphs on the production costs and the reluctance of local populations (that seems to be fuelled by right-wing parties). Chapter 5 is specifically about solar photovoltaïcs, deemed to be now cheaper than fossil fuels inmost countries. And substituting for deficient or inexistent large scale energy grids in some countries. And Chapter 6 deals (briefly) with other low-carbon technologies, like tidal dams (mentioning the 1966 La Rance dam near Mont Saint-Michel, Normandy, we would visit now and then when I was a kid!), wave turbines, nuclear energy, and geothermal power (both heating and providing electricity). Chapter 7 discusses renewable electricity issues with energy storage (batteries and pumped hydro storage), since most solutions cannot be fired at will. The book  addresses neither the loss in carrying electricity over long distances (as suggested p103 between Morocco and Europe, or Australia and Singapore), nor the hacking risks impacting large electricity grids. Chapter 8 switches to decarbonasing heat and transport, where heat pumps and electric vehicles are the most promising venues. Chapter 9 concludes by a more political discussion of the transition to renewable, pointing out the Chinese leadership in switching to solar and wind capacities. And the brake put on the transition by international crisis such as the Russian invasion of Ukraine that keep subsidies on fuel consumption. And make European countries divesting from this transition to invest in military budgets.

While the book manages a proper introduction to renewable energy and stays up-to-date with the current developments, I find it a bit overly optimistic on the prospect of achieving COP goals and carbon neutrality. Beyond the geostrategic issues briefly mentioned in the concluding chapter, there is no mention made of the exploding energy consumption of AIs and of the limited investments of AI companies into renewable energies… Reducing energy demand does not even occupy one page of the book (p127). Similarly, I find too little discussion of the political and human aspects of using renewables, eg photovoltaïcs and batteries, which resurfaced in the recent Chinese blockade on rare earths or coverages (as in Nature, 04 Nov 2025) on the extreme hardship of extracting minerals. Contrary to those (aspects) for massive dams affecting the local populations and in the dispute between countries or States. And also little on the environmental costs of producing and recycling both solar and wind farms, in contrast with hydroelectricity. Surprisingly, nuclear energy is evacuated in one paragraph in Chapter 2, on safety arguments. If reappearing in Chapter 6 with further concerns about the overall cost of nuclear energy.

[The usual disclaimer applies, namely that this bicephalic review is likely to appear later in CHANCE, in my book reviews column.]

scalable Monte Carlo for Bayesian learning [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , on September 26, 2025 by xi'an

This book by Paul Fearnhead, Christopher Nemeth, Chris Oates, and Chris Sherlock is part of the IMS Monograph series. And published by Cambridge University Press. It covers most recent developments in MCMC methods, namely stochastic gradient MCMC (Chap. 3), non-reversible MCMC (Chap. 4), continuous-time MCMC (Chap. 5), and assessing and improving MCMC (Chap. 6). I find the book remarkable in its attention to rigour and clarity, without falling into overly technical derivations. It is perfectly suited for a graduate course to students with a solid mathematical background. In short, had I considered a new edition of our Monte Carlo Statistical Methods book to incorporate these advances, I could not done such a good job!

The first chapter provides a quick refresher of the background, from Monte Carlo principles, to Markov chains, SDEs, and the kernel “trick” (which requires a dozen pages of exposition). Nonetheless, it contains side remarks of true interest, including some suggestions I had not previously seen, as for instance an unusual introduction of the HMC algorithm as an underdamped Langevin diffusion. Chapter 2 prolongates this recap by covering reversible MCMC algorithms and the attached optimal scalings. This is done in a particularly friendly presentation that I intend to use in my own course. The HMC section is probably the best coverage I have seen on the topic, including most naturally the leapfrog steps.

Chapter 3 gets into stochastic gradient MCMC as an approximate MCMC, with nice arguments and formal convergence bounds. Again quite efficiently, if focussing almost solely on Gaussian settings (but including a neural network example). Similarly, Chapter 4 provides intuitive (if informal) arguments on the worth of non-reversible algorithms that are well-suited to a textbook of this level. This chapter introduces a PDMP sampler like the discrete bouncy particle sampler.

Chapter 5 is a (nicely) monstrous coverage of continuous time MCMC samplers that reaches very recent advances on PDMPs. The focus is on expressing them as limits, in order to derive mixing rates without extreme mathematical steps. (The chapter even includes a mention to the coordinate sampler that my PhD student Wu Changye derived in 2018!) Again a chapter I plan to use when teaching MCM methods, if possibly skipping some of the 66 pages.

Chapter 6 completes the monograph with a presentation of convergence assessment tools and diagnostics, exploiting the kernel trick, as well as convergence bounds that reflect very recent research in that domain. The conclusive section on optimal weights and optimal thinning will presumably be new to most readers. (Making me wonder if a link can be found with our importance Markov chain construct.)

[Disclaimer about potential self-plagiarism as usual: this post or an edited version will eventually appear in my Books Review section in CHANCE.]

A modern introduction to probability and statistics [book review]

Posted in Books, R, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , on July 12, 2025 by xi'an

In the plane to Bengaluru, I read through the book A modern introduction to probability and statistics, by Graham Upton—whose Measuring Animal Abundance I reviewed for CHANCE a while ago—, which is based on the earlier Understanding Statistics, written jointly with Ian Cook. (Not to be confused with A modern introduction to probability and statistics by Dekking et al.) The subtitle is understanding statistical principles in the computer age. Sorry, in the age of the computer. While the cover is most pleasant (and modern), as noticed by an AF flight attendant, the contents are very very standard and could have been written decades ago since the main concession to “the” computer age is the inclusion of a few R commands at the end of most chapters. There are even a few distribution tables here and there (in case “the” computer is not available). But there is no other connection with computational statistics or statistical computing.

The classicism of the contents and the intended audience mean there is little therein on which to either object or criticise. The mixture of elementary probability and basic statistics in a single textbook always feels awkward to me and I think I would have trouble teaching solely from this material. Apart from the glaring typo on the variance of the sum of two correlated random variables on page 87, missing the factor 2 in front of the covariance, while correct(ed) p97 (and the inevitable “the the” typo spotted once). My main criticisms are on the potential confusion between samples and populations in the early chapters, when some statistics are used as motivational examples, as for instance in a (hidden) Monte Carlo stabilisation to the limiting values (p57), way before the Law of Large Numbers is introduced,, the variable mileage in mathematical rigour (while being uncertain that first year students can handle integrals and derivatives), the textbook examples, and the amount of the book contents spent on descriptive statistics and even more on the “classical” tests, with no critical perspective on using point nulls or p-values. The book concludes with a four page (benevolent) chapter on Bayesian statistics that is superfluous imho, or even counterproductive since my experience with a rushed introduction to Bayesian principles almost always result in a rejection of said principles. Plus, the illustration with the coin tossing is not particularly helpful since Andrew maintains that one can load a die, but cannot bias a coin. (A similar reservation on the half-page 289 coverage on pseudo-random generation and Monte Carlo principles for computing p-values.)

Minor (mostly idiosyncratic) remarks follow: CLT prior to LLN,   n-1 in sample sd, little to no model criticism (ntbcf goodness of fit), missing an opportunity when mentioning the varying probability of a day being a birthday (p31) in contrast with BDA cover story, and another opportunity to cite the 2024 Ig Nobel Prize for coin tossing around the LLN, an unclear definition for random variables( p53) and a potentially confusing introduction of Poisson distributions through a informal reference to Poisson processes (and no reason why the years of accession of the kings of Sussex and England till Guillaume—making a return on p178 with the Domesday Book—in 1066 should follow such a process as suggested in Figure 3.5), a surprising definition of the constant e as the special case of exp(x) when x=1 and its series expansion (p70), omitting proofs on laws of sums of iid rv’s by introducing moment generating functions rather late, another obscure reference to a 16th German treatise on surveying as a precursor of the CLT (p131), a proof for the normalising constant of the Normal density that will most likely escape most first year students, a introduction of the t, F, and χ² distributions with no mention of their respective densities (pp141-147), never defining a joint Normal distribution density, insisting on unbiasedness without noting that maximum likelihood—with a strange motivation that it “makes the next sample of n observations most likely to resemble the data in the current sample (p228)—estimators are almost always biased, an abundance of footnotes that may prove of little interest for the youngest readers.

[Disclaimer about potential self-plagiarism as usual: this post or an edited version will eventually appear in my Books Review section in CHANCE.]

AI Narratives [book review]

Posted in Books with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on June 13, 2025 by xi'an

AI Narratives: A history of imaginative thinking about intelligent machines is a 2020 collective book edited by Stephen Cave, Kanta Dihal, and Sarah Dillon, with about twenty contributing authors, through a series of 16 chapters on the relation between culture (literature, films) and our societal approach to AI, with varying perspectives, some overlap between chapters, a wide range of extrapolation, especially in relation with the oldest books (like Homer’s) and quotes from both books I enjoyed and books I had not heard of, to add to my to-red pile, like Roderick. Predominant place of Blade Runner, of Ĉapek’s Rossum’s Universal Robots, not only for introducing the term, and of Asimov, but also a detailed analysis of the great, thought-provoking, Ann Leckie’s Ancillary Justice. With a realization that Gibson’s Neuromancer had aged quite a lot… For the modern times, very little outside the US-UK realm, as for instance no mention made of Jules Verne or René Barjavel, although Zola makes an appearance when describing workers as automats, neither of the Germanic literature (from the Grimm Brothers onwards), nor of the non-negligible USSR science fiction production, nor yet of the Chinese input, like The three body problem. Other books that could have made it: the Murderbot diaries, and The Alchemy Wars trilogy, for digging into the blurry border between humans and AIs; A Memory Called Empire, for a clever approach to mind uploading, as well as the masterly Never let me go by Ishiguro, both dealing with a future where copies of humans The Amazing Adventures of Kavalier & Clay for the golem, the short story An Unatural Life, for its highly original take on the legal rights of humanoid robots, A Psalm for the Wild-Built because of… tea and Zen monk, obviously!

Among things I learned from AI Narratives, the story of Alan Turing (figuratively) mansplaining Ada Lovelace on her pronouncement that machines cannot be intelligent since they deliver what they are coded for. The mention of automatos in the Illiad. A mention of 1625 Gabriel Naudé Apologie &tc. that excludes magic as irrational, as well as introducing the term androide. The trivia that the creator of the (fraudulent) automaton chess player, Kempelen, also produced in 1791 an authentic if brainless speaking machine. The realisation that E.M Forster also wrote a futuristic novel, The Machine Stops (as well as the precursor of the symbolists, Villiers de l’Isle Adam, with L’Ève Future). Another trivia that cyberpunk first appeared in a 1983 short story by Bruce Bethke. A discussion of the elaborate Culture constructed by Iain Banks, albeit through volumes in the series I hade not read, along with the concept of OCP for outside context problem, akin to Taleb’s Black Swan. The least interesting (and final) chapter in the book is paradoxically the closest to data analysis, when Recchia runs a rather low-tech assessment of a subtitle dataset in relation with AI, if referring to Tufte’s rule of the baselin in the notes (p405).

Definitely enjoyable book, then, even though I mostly skimmed through the chapters during a day-trip to Lille on a Bank (!) holiday. Reading through it made me muse rather belatedly of a parallel between AI scare or adoration, and the non-AI societies, where individuals are (also) part of a structure large enough to miss the larger picture. From building pyramids to being part of the global economy. (This followed mostly from my surprise in seeing Dickens, Trollope, and Zola included in the discussion.)

[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE.]