Archive for prior selection

veniSBA²

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , on July 4, 2024 by xi'an

After another morning cycle of 2Xing Porte della Libertà (under a light and pleasant rain) and swimming in Sant’ Alviso (in too warm a water), I did not make it for the beginning of the Bayesian deep learning session, breakfast oblige!, and cumulated with different percolation events (ie, meeting friend after friend on my way to the classroom), I could not get enough of the session to report anything even barely useful!

As I did not rush fast enough to Andrew’s Foundation lecture (another sequence of percolations!), I had to stand in the back of the packed main amphitheatre (and former sorting hall of the Venice slaughterhouse!), Guido Cazzavillan’s Aula Magna, while he talked a fresco about some holes in Bayesian data analysis (the analysis, not the book!), those being [verbatim]

  1. the usual rules of conditional probability fail in the quantum realm,
  2. flat or weak priors lead to terrible inferences about things we care about,
  3. subjective priors are incoherent,
  4. Bayesian decision picks the wrong model,
  5. Bayes factors fail in the presence of flat or weak priors,
  6. for Cantorian reasons we need to check our models, but this destroys the coherence of Bayesian inference.

After lunch, I attended the (mostly sequential) simulation based inference (renamed from ABC!) session with a composite likelihood proposal by Lorenzo Rimella, that uses marginals to approximate the likelihood of a hidden Markov SIS epidemic model by composite likelihood towards getting more efficient if inexact versions. Then [1WABC webinar co-organiser] Umberto Picchini on surrogates for likelihood and posterior functions, with sequential improvements (w/o ABC and w/o neural networks). Called “Sequential mixture posterior and likelihood estimation”, using mixtures of experts when the weights are functions of the observed or simulated y. With adapting the number of components in the mixture. Comparing favourably with normalising flows. And Wentao Li on correcting by ABC for composite likelihood as in Ruli et al. (2016). Where a posterior distribution given composite scores (seen as [summary] statistics) is employed but requires a convergent estimator of the unknown parameter.

 No congratulation today to our PhD student who managed to fall in a canal (but survived)..!

Philosophies, Puzzles and Paradoxes [book review]

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on May 25, 2024 by xi'an

Yudi Pawitan and Youngjo Lee have written a book that recently caught my attention within the CRC Press list of new publications. Because philosophy, puzzles, and paradoxes are definitely of interest to me (as shown by numerous entries in the ‘Og!). The subtitle of said book is A Statistician’s Search for Truth.

Reviews of the book are already available, with for instance Andrew Gelman stating that he disagrees “with much of this book, but it’s an entertaining and thought-provoking introduction to some challenging questions” or Stephen Senn starting the foreword with “This is a remarkable book: wide-ranging, ambitious, challenging and profound but also intriguing, fascinating and original.” (Senn is also cited within the book for his discussion of our revisit of Harold Jeffreys’ Theory of Probability.) Nice cover as well (albeit I could not trace the origin of it, inside or outside the book.)

The book is made of three parts, one on the philosophical approaches to truth, scientific discovery, deduction, and induction, a second one on probability theories, with philosophical motivations, Bayesian inference, and likelihood-based inference, and a third section on paradoxes. Given that both authors are senior authors who have contributed to likelihood inference throughout their career, incl. the books In All Likelihood and Generalized Linear Models with Random Effects, the likelihood approach is somewhat privileged against other statistical resolutions towards the resolution of the paradoxes, with a defence of confidence distributions and a chapter on epistemic confidence that mostly stems from recent papers by the authors, like Pawitan et al.  (2023) and Lee and Lee (2023). I find the discussion therein somewhat unclear, esp. because the same notation Pr(.) is employed for different probability notions.

“Epistemic confidence is the objective measure of uncertainty that’s attached to single events, where the objectivity is based on a consensus of rational minds.” (p.197)

The philosophy part is following the (European) Enlightenment in producing more and more involved discussions on reason, knowledge and scientific discovery. This exploration is an easy read, as it does not delve particularly deeply in the arguments of Kant, Hume, or Popper. With the apparently unescapable mention of Gödel’s incompleteness theorem, including a sausage citation from Poincaré that reminded of that strip from Tintin in America:which, most probably, he would have applied to Ais! Several sections about pseudo-rational attempts to demonstrate the existence of Dog could have been skipped as well.

The part of probability already considers paradoxes which, like the subsequent ones are mostly the consequence of using natural (and hence ambiguous) languages instead of mathematical descriptions—incl. the statement of the Likelihood Principle. It also discusses Keynes’ logical (or imprecise) probabilities, briefly if appropriately given the pessimistic views of young Keynes on the assessment of the probability of an event. Savage is privileged enough to enjoy an entire chapter discussing his 1950’s axioms leading to the existence of a (subjective)  prior on “the states of the world”. This is followed by a chapter on Inverse probability (aka Bayesian statistics), where the authors consider Bayes’ 1763 Essay to have stayed mostly unnoticed till  the beginning of the 20th Century, which sounds a somewhat subjective judgement. (And as uncovered by Steve Stiegler, the original title of the Essay was indeed intended as a reply to Hume.) A further if short chapter is dedicated to the search for the prior distribution. Which thus gives the misguided impression that there should exist such a thing, rather than acknowledging that Bayesian statements are relative to the prior measure. The remainder of the discussion on invariant and reference priors is however mostly standard. Except when falling for the marginalisation paradox when stating that a product of improper priors implies independence on p.144.

The paradoxes examined in the final part are Allais’ (an alumni of Lycée Lakanal!), and Ellsberg’s, avatars of the Saint Petersburg paradox and referring to failing to adhere to rational decision-making and not in the least to statistics. Conjunction and inclusion “fallacious fallacies”, which are central to Kahneman’s Thinking fast and slow bestseller, and attributed to reasoning in terms of likelihood rather than of probability (without accounting for multiple testing on p.228). A whole if short chapter on the Monty Hall and three prisoners paradoxes, another predictable occurrence in a book on reasoning paradoxes. Again mostly a matter of poor wording, plus relying on the choice of an underlying probability model, for which the authors again follow a likelihood approach, the number of the prize door or of the freed prisoner being the parameter. Kyburg’s (very weak) lottery paradox and related forensic paradoxes, concluding with the rejection of judgements based solely on probability reasoning. Hempel’s paradox of the ravens, a priori unrelated with statistical evidence, but turned into one by squeezing in some sampling models. Finishing with the (envelope) exchange paradox, where the authors refuse to put a prior on the unknown parameter but end up with a solution equivalent to adopting a Jeffreys prior.

In conclusion, this attempt at connecting statistical inference and philosophy, probability concepts and rational decision making, paradoxes and modelling, within a single book is academically sound and overall enjoyable, if not outstanding or remarkable as it does not constitute a radical move away from existing analyses of those classical paradoxes. Furthermore, I find the paradoxes overwhelmingly distant from genuine statistical settings and involving a rather stretched notion of data. Still, methinks I will keep this book in my bookcase, rather than leaving it for the taking in the department coffee room!

As I was completing the book and getting towards writing this book review, I also noticed a two page blurb in Significance (May 2024 issue) written by the authors on their book. (which happens rather frequently with this magazine). Unsurprisingly, the contents provd mostly extracted from the preface and introduction With a nice ravens picture (in conjunction with the raven paradox).

[Disclaimer about potential self-plagiarism: this post or an edited version may eventually appear in my Books Review section in CHANCE.]

Bertrand’s paradox [re]solved?

Posted in Books, pictures, Statistics, Travel with tags , , , , , , , , , , , on September 29, 2023 by xi'an

On the plane back from Vancouver, I read Bertrand’s Paradox Resolution and Its Implications for the Bing–Fisher Problem by Richard A. Chechile [who had pointed out his paper to me] In this paper, Chechile considers the Bayesian connections/sequences of Betrand’s paradox, as he sees it Bertrand’s different solutions/paradox to be

“designed to illustrate his dissatisfaction with the Bayes and Laplace use of a probability distribution to represent an unknown parameter that can have any continuous value”

and proposes to “resolve” this paradox, which imho is neither a paradox nor in need of a resolution!, as I see it more like a reflection on the importance of sigma algebras and measure theory. The uniform distribution (behind the “random” chord) is not a uniquely specified concept, just like the maximum entropy distribution is relative to the dominating measure. When arguing that

“Such a definition [based on any possible distribution of a stochastic chord] would yield a random variable, but this weak sense of the word random is not satisfactory, because there is an infinite number of stochastic processes that can be defined to yield a probability distribution of chord lengths.”

the author is simply restating that infinite collection of dominating measures.  But imho he is somewhat missing this point when defining Shannon`s entropy by resorting to a discrete version. And when adopting a uniform measure on the chord as a reference (Section 3.2, on The Importance of a Dominant Metric Representation). While the probability P(L>1) is invariant under any increasing transform of L (and 1)… This amounts to arguing for a favourite parameterisation in constructing  a reference prior (Section 4, where Jeffreys prior is also dismissed for not being at maximum entropy). The ensuing discussion as to why the three solutions of Bertrand’s are not valid (Section 2.2) is thus most curious to me since they all are implementable/practical ways of producing stochastic chords. I find it rather amusing that one returns to the quest for the ideal priori distribution Bayesians were so fiercely debating at the turn of the previous century. And non-Bayesians were all too happy to exploit when arguing against this approach.

statistical modeling with R [book review]

Posted in Books, Statistics with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on June 10, 2023 by xi'an

Statistical Modeling with R (A dual frequentist and Bayesian approach for life scientists) is a recent book written by Pablo Inchausti, from Uruguay. In a highly personal and congenial style (witness the preface), with references to (fiction) books that enticed me to buy them. The book was sent to me by the JASA book editor for review and I went through the whole of it during my flight back from Jeddah. [Disclaimer about potential self-plagiarism: this post or a likely edited version of it will eventually appear in JASA. If not CHANCE, for once.]

The very first sentence (after the preface) quotes my late friend Steve Fienberg, which is definitely starting on the right foot. The exposition of the motivations for writing the book is quite convincing, with more emphasis than usual put on the notion and limitations of modeling. The discourse is overall inspirational and contains many relevant remarks and links that make it worth reading it as a whole. While heavily connected with a few R packages like fitdist, fitistrplus, brms (a  front for Stan), glm, glmer, the book is wisely bypassing the perilous reef of recalling R bases. Similarly for the foundations of probability and statistics. While lacking in formal definitions, in my opinion, it reads well enough to somehow compensate for this very lack. I also appreciate the coherent and throughout continuation of the parallel description of Bayesian and non-Bayesian analyses, an attempt that often too often quickly disappear in other books. (As an aside, note that hardly anyone claims to be a frequentist, except maybe Deborah Mayo.) A new model is almost invariably backed by a new dataset, if a few being somewhat inappropriate as in the mammal sleep patterns of Chapter 5. Or in Fig. 6.1.

Given that the main motivation for the book (when compared with references like BDA) is heavily towards the practical implementation of statistical modelling via R packages, it is inevitable that a large fraction of Statistical Modeling with R is spent on the analysis of R outputs, even though it sometimes feels a wee bit too heavy for yours truly.  The R screen-copies are however produced in moderate quantity and size, even though the variations in typography/fonts (at least on my copy?!) may prove confusing. Obviously the high (explosive?) distinction between regression models may eventually prove challenging for the novice reader. The specific issue of prior input (or “defining priors”) is briefly addressed in a non-chapter (p.323), although mentions are made throughout preceding chapters. I note the nice appearance of hierarchical models and experimental designs towards the end, but would have appreciated some discussions on missing topics such as time series, causality, connections with machine learning, non-parametrics, model misspecification. As an aside, I appreciated being reminded about the apocryphal nature of Ockham’s much cited quote “Pluralitas non est ponenda sine necessitate“.

Typo Jeffries found in Fig. 2.1, along with a rather sketchy representation of the history of both frequentist and Bayesian statistics. And Jon Wakefield’s book (with related purpose of presenting both versions of parametric inference) was mistakenly entered as Wakenfield’s in the bibliography file. Some repetitions occur. I do not like the use of the equivalence symbol ≈ for proportionality. And I found two occurrences of the unavoidable “the the” typo (p.174 and p.422). I also had trouble with some sentences like “long-run, hypothetical distribution of parameter estimates known as the sampling distribution” (p.27), “maximum likelihood estimates [being] sufficient” (p.28), “Jeffreys’ (1939) conjugate priors” [which were introduced by Raiffa and Schlaifer] (p.35), “A posteriori tests in frequentist models” (p.130), “exponential families [having] limited practical implications for non-statisticians” (p.190), “choice of priors being correct” (p.339), or calling MCMC sample terms “estimates” (p.42), and issues with some repetitions, missing indices for acronyms, packages, datasets, but did not bemoan the lack homework sections (beyond suggesting new datasets for analysis).

A problematic MCMC entry is found when calibrating the choice of the Metropolis-Hastings proposal towards avoiding negative values “that will generate an error when calculating the log-likelihood” (p.43) since it suggests proposed values should not exceed the support of the posterior (and indicates a poor coding of the log-likelihood!). I also find the motivation for the full conditional decomposition behind the Gibbs sampler (p.47) unnecessarily confusing. (And automatically having a Metropolis-Hastings step within Gibbs as on Fig. 3.9 brings another magnitude of confusion.) The Bayes factor section is very terse. The derivation of the Kullback-Leibler representation (7.3) as an expected log likelihood ratio seems to be missing a reference measure. Of course, seeing a detailed coverage of DIC (Section 7.4) did not suit me either, even though the issue with mixtures was alluded to (with no detail whatsoever). The Nelder presentation of the generalised linear models felt somewhat antiquated, since the addition of the scale factor a(φ) sounds over-parameterized.

But those are minor quibble in relation to a book that should attract curious minds of various background knowledge and expertise in statistics, as well as work nicely to support an enthusiastic teacher of statistical modelling. I thus recommend this book most enthusiastically.

reproducibility check [Nature]

Posted in Statistics with tags , , , , , , , , on September 1, 2021 by xi'an

While reading the Nature article Swarm Learning, by Warnat-Herresthal et [many] al., which goes beyond federated learning by removing the need for a central coordinator, [if resorting to naïve averaging of the neural network parameters] I came across this reporting summary on the statistics checks made by the authors. With a specific box on Bayesian analysis and MCMC implementation!