I spotted this title in the new arXiv postings on Monday. When Is Generalized Bayes Bayesian? A Decision-Theoretic Characterization of Loss-Based Updating by Kenichiro McAlinn & Kōsaku Takanashi is discussing decision-theoretic consequences of generalized Bayes approaches based on losses and show that decisions based on a loss-based posterior coincides with those of ordinary Bayes if and only if the loss is essentially a negative log-likelihood (leading to a belief posterior). This is not very surprising in that, otherwise, there is no Bayesian update delivering the generalised Bayes pseudo-posteriors (which can be traced back to a 2007 result of Catoni). The authors also demonstrate that generalized marginal likelihoods are not delivering evidence for decision posteriors, and thus that Bayes factors are not well-defined in this context, which reminds me of our warning for ABC model choice. However, the reason here is much more mundane, as it is due to the decision posterior failing to identify the normalising constant Z(x). Outside belief posteriors. The paper concludes with a coherence book, which is a table reproduced above.
Archive for Bayes factors
When Is Generalized Bayes Bayesian?
Posted in Books, Statistics, University life with tags ABC, Bayes factors, Bayesian model choice, coherence, decision theory, generalised Bayesian inference, loss functions, marginal likelihood, normalising constant on February 13, 2026 by xi'ane-values in Chennai
Posted in Books, pictures, Running, Statistics, Travel, University life with tags Abraham Wald, admissibility, Bayes factors, Bayesian hypothesis testing, Benjamini, BIRSCMI, Chennai, Chennai Mathematical Institute, complete class theorems, confidence sequence, Dickey-Savage ratio, e-values, empirical Bayes methods, FDRs, Hochberg, Neyman-Pearson tests, p-values, sequential testing, south Indian cuisine, Tamil Nadu on July 23, 2025 by xi'an
To recap, I thus attended the BIRS-CMI workshop 25w5482 at the Chennai Mathematical Institute, Navalur, Tamil Nadu, in early July, for being intrigued by the developments around the concept. And enjoyed the week, from partaking in the company of friendly and enthusiastic academics to the exposure of new views and concepts, mostly remote from mine’s. Recall that an e-value attached to an hypothesis H described as a collection of distributions is a non-negative random variable E with expectation less than 1 for E~Q and all Q ∈ H. When a stopping rule is involved, the e-value is extended into an e-process. (Beyond Aaditya Ramdas’ E-book, Ruodu Wang also wrote a “tiny” review.) Aaditya Ramdas recalled in his introduction of the workshop that e-values are fundamentally equivalent to p-values and confidence intervals. And that a confidence sequence is a sequence of confidence intervals that contains the true value for all time steps t’s with a probability of at least 1-α.
The talks reflected a general belief in α levels and in Neyman-Pearsonian likelihood ratio optimality in simple vs simple settings, considering extension for sequential analysis settings, anytime inference, universality under general alternatives, and connections with FDRs, incl. Benjamini & Hochberg solution, but pointed out a lack of middle ground between frequentists and Bayesians.
“e-values have a clear interpretation in terms of betting and are closely related to likelihood ratios and other Bayes factor. At the same time, e–values do not require prior distributions conditional on the null and alternative hypotheses”

Although David R. Bickel attempted a Bayesian version, using a marginal likelihood ratio within betting settings, that is an incoming American Statistician paper. I may have being missing some aspects due to a lack of sleep the night before (!), but I find the attempt resulting in a fairly unusual vision of Bayesian testing as either not depending on any parameter or on the opposite using a family of priors. I did not understand either the “criticism” that the predictive depends on the prior and felt that this representation was bending in a rather onsiderable way the Bayesian perspective towards achieving a certain degree of agreement with p– and e-value notions, to conclude that the Bayes factor is an e-value. (As an aside, this may be the first paper that cited our critical review of Aitkin! Similarly, Shubhada Agrawal mentioned Roger Farrell in his talk, with whom we wrote a complete class Annals paper in the late 1980’s.) Nikos Ignatiadis also explored Empirical Bayes e-values, while Ben Chugg gave a presentation (constrained) admissibility, albeit under type-I error constraints that makes Bayes infeasible and using Neyman-Pearsonian loss functions. On the last day, Peter Grünwald tried for some BFF cohesion with openings on e-posteriors, treating hypothesis testing losses symmetrically, defining it as an inverse of e-values but incorporating pseudo-posteriors of many flavours like confidence, inferential, and fiducial distributions. He also mentioned a Savage-Dickey version while using an arbitrary prior, which is also an e-value, but with upper & lower meanings, again with measure issues
Given the hosting of the workshop in the Chennai Mathematical Institute, which is quite far from the centre of town (much closer to Mahabalipuram!), I did not visit Chennai but enjoyed the South Indian cuisine (albeit missing some fierceness in the spices!) and local fruits from street stands, if being sorry I could not find cocoa pods from nearby Kerala.
Bayesian Inference: Theory, Methods, Computations [book review]
Posted in Statistics with tags ABC, ABC-MCMC, Bayes factors, Bayesian decision theory, Bayesian inference, Bayesian testing, Bayesian textbook, BIC, book review, capture-recapture, CHANCE, Chapman & Hall, CRC Press, DIC, Jeffreys priors, lizards, statistical inference, subjective versus objective Bayes, The Bayesian Choice, toe clipping, Uppsala University, variational Bayes methods on November 12, 2024 by xi'an
Bayesian Inference: Theory, Methods, Computations by Silvelyn Zwanzig and Rauf Ahmad, both from Uppsala University, is a recent book published by Chapman & Hall / CRC Press. About 300p long (plus appendices), it covers the core aspects of Bayesian inference, namely the decision theoretic motivations, its asymptotic validation, the specifics of estimation and testing, and the computational approximations (MC, MCMC, ABC, VB), with entries on prior specification and Normal linear models. And some R codes. It is (and feels like) constructed from Master and PhD courses (at Uppsala University), with a rigorous mathematical presentation and many examples, some related to biostatistics. Drawings from the first author’s daughter are included in most chapters, to this reviewer’s bemusement. From a further personal viewpoint, the book also reads rather close to my (Bayesian) choice of a Bayesian textbook, which proves rather accurate since several chapters are inspired by my own Bayesian Choice. as acknowledged therein. As well as by the more recent Statistical Decision Theory: Estimation, Testing, and Selection by Liese & Miescke (2008) and Introduction to the Theory of Statistical Inference by Liero & Zwanzig (2011). Witness, for instance, an example of prior construction for capture-recapture experiments on lizards as analysed by my PhD student Dupuis (1995) [with a curious switch to the authors on p.263] and also included in The Bayesian Choice (with drawing 2.9 incorrect in that the lizards there have marks on their backs, instead of the code adopted by the ecologists, namely cutting one specific phalange for each capture).
Other minor quandaries: The usual issue of quoting the wrong edition for creating a method, as when citing Jeffreys (1946) for inventing non-informative priors [p.53], failing to point out the parameterisation invariance of intrinsic losses [p.95]considering that Bayes factors are only relevant for obtaining evidence against the null hypothesis [p.216], recommending BIC and DIC (!) [pp.232-6], advocating sampling importance resampling (SIR) for approximate sampling from the target (omitting infinite variance issues) [p.253], defining annealing as using “several trial distributions” [p.261], a mistake in ABC-MCMC [p.274] since the case when the simulated data is too far from the actual data should lead to a repetition rather than a pure rejection.
All in all, a reasonable textbook with some recent input, but still lacking in originality, if I may subjectively say so.
[Disclaimer about potential self-plagiarism: this post or an edited version of it could possibly appear in my Books Review section in CHANCE.]
statistical modeling with R [book review]
Posted in Books, Statistics with tags AIC, Bayes factors, Bayesian Analysis, Bayesian data analysis, book review, brms, CHANCE, conjugate priors, Deborah Mayo, DIC, fitdist, fitistrplus, fonts, frequentist inference, Gibbs sampling, glm, glmer, JASA, Jeddah, Jeffreys priors, Journal of the American Statistical Association, machine learning, MCMC, Metropolis-Hastings algorithm, model misspecification, non-parametrics, Ockham's razor, OUP, Oxford University Press, packages, plagiarism, prior selection, R, STAN, Statistical Modeling, Steve Fienberg, support, Uruguay, WAIC on June 10, 2023 by xi'anStatistical Modeling with R (A dual frequentist and Bayesian approach for life scientists) is a recent book written by Pablo Inchausti, from Uruguay. In a highly personal and congenial style (witness the preface), with references to (fiction) books that enticed me to buy them. The book was sent to me by the JASA book editor for review and I went through the whole of it during my flight back from Jeddah. [Disclaimer about potential self-plagiarism: this post or a likely edited version of it will eventually appear in JASA. If not CHANCE, for once.]
The very first sentence (after the preface) quotes my late friend Steve Fienberg, which is definitely starting on the right foot. The exposition of the motivations for writing the book is quite convincing, with more emphasis than usual put on the notion and limitations of modeling. The discourse is overall inspirational and contains many relevant remarks and links that make it worth reading it as a whole. While heavily connected with a few R packages like fitdist, fitistrplus, brms (a front for Stan), glm, glmer, the book is wisely bypassing the perilous reef of recalling R bases. Similarly for the foundations of probability and statistics. While lacking in formal definitions, in my opinion, it reads well enough to somehow compensate for this very lack. I also appreciate the coherent and throughout continuation of the parallel description of Bayesian and non-Bayesian analyses, an attempt that often too often quickly disappear in other books. (As an aside, note that hardly anyone claims to be a frequentist, except maybe Deborah Mayo.) A new model is almost invariably backed by a new dataset, if a few being somewhat inappropriate as in the mammal sleep patterns of Chapter 5. Or in Fig. 6.1.
Given that the main motivation for the book (when compared with references like BDA) is heavily towards the practical implementation of statistical modelling via R packages, it is inevitable that a large fraction of Statistical Modeling with R is spent on the analysis of R outputs, even though it sometimes feels a wee bit too heavy for yours truly. The R screen-copies are however produced in moderate quantity and size, even though the variations in typography/fonts (at least on my copy?!) may prove confusing. Obviously the high (explosive?) distinction between regression models may eventually prove challenging for the novice reader. The specific issue of prior input (or “defining priors”) is briefly addressed in a non-chapter (p.323), although mentions are made throughout preceding chapters. I note the nice appearance of hierarchical models and experimental designs towards the end, but would have appreciated some discussions on missing topics such as time series, causality, connections with machine learning, non-parametrics, model misspecification. As an aside, I appreciated being reminded about the apocryphal nature of Ockham’s much cited quote “Pluralitas non est ponenda sine necessitate“.
Typo Jeffries found in Fig. 2.1, along with a rather sketchy representation of the history of both frequentist and Bayesian statistics. And Jon Wakefield’s book (with related purpose of presenting both versions of parametric inference) was mistakenly entered as Wakenfield’s in the bibliography file. Some repetitions occur. I do not like the use of the equivalence symbol ≈ for proportionality. And I found two occurrences of the unavoidable “the the” typo (p.174 and p.422). I also had trouble with some sentences like “long-run, hypothetical distribution of parameter estimates known as the sampling distribution” (p.27), “maximum likelihood estimates [being] sufficient” (p.28), “Jeffreys’ (1939) conjugate priors” [which were introduced by Raiffa and Schlaifer] (p.35), “A posteriori tests in frequentist models” (p.130), “exponential families [having] limited practical implications for non-statisticians” (p.190), “choice of priors being correct” (p.339), or calling MCMC sample terms “estimates” (p.42), and issues with some repetitions, missing indices for acronyms, packages, datasets, but did not bemoan the lack homework sections (beyond suggesting new datasets for analysis).
A problematic MCMC entry is found when calibrating the choice of the Metropolis-Hastings proposal towards avoiding negative values “that will generate an error when calculating the log-likelihood” (p.43) since it suggests proposed values should not exceed the support of the posterior (and indicates a poor coding of the log-likelihood!). I also find the motivation for the full conditional decomposition behind the Gibbs sampler (p.47) unnecessarily confusing. (And automatically having a Metropolis-Hastings step within Gibbs as on Fig. 3.9 brings another magnitude of confusion.) The Bayes factor section is very terse. The derivation of the Kullback-Leibler representation (7.3) as an expected log likelihood ratio seems to be missing a reference measure. Of course, seeing a detailed coverage of DIC (Section 7.4) did not suit me either, even though the issue with mixtures was alluded to (with no detail whatsoever). The Nelder presentation of the generalised linear models felt somewhat antiquated, since the addition of the scale factor a(φ) sounds over-parameterized.
But those are minor quibble in relation to a book that should attract curious minds of various background knowledge and expertise in statistics, as well as work nicely to support an enthusiastic teacher of statistical modelling. I thus recommend this book most enthusiastically.
van Dantzig seminar
Posted in pictures, Statistics, Travel, University life with tags Amsterdam, Bayes factors, Centrum Wiskunde & Informatica, Chib's approximation, CWI, David van Dantzig, evidence, mixtures of distributions, seminar, Thalys, the Netherlands, Theory of Collective Phenomena, Van Dantzig Seminar on June 3, 2023 by xi'an