Archive for The Likelihood Principle

Seminal ideas and controversies in Statistics [book review]

Posted in Books, Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on May 24, 2025 by xi'an

CRC Press sent CHANCE this book for review. Since the topic was of clear interest to me, with an author who significantly contributed to the field—my only recollection meeting Roderick Little was during the Australian Statistical Conference in Adelaïde, in 2012, at the start of my Oz 2012 Tour!—, I took the opportunity of the nearest weekend to browse through Seminal ideas and controversies in Statistics. I like very much the idea of selecting a dozen key papers in the history of Statistics and of discussing why. In fact, this reminded me of my classics seminar, which lasted the few years I was 100% in charge of the Master program in Dauphine (and which I hope I could restart!). Checking the list of the papers I then suggested my students, I see some overlap with 9 papers out of the 15 groups. (I also remember Steve Fienberg making suggestions for that list, while he was spending a sabbatical in Paris at CREST.) Given that community of focus and purpose, and contrary to my wont, I have really very little of substance to criticize or wish about the book. The less when reading the following

“On a personal note, I met Yates [author of a 1984 paper on tests for 2×2 contingency tables discussing the relevance of conditioning on one or both margins], a charming man, when I was a young graduate student who knew next to nothing about statistics; we discussed the joys of traversing the Cuillin Ridge in Skye.”

since completing that ridge remains high in my mountain-climbing bucket-list! (Possibly next year, since we are running an ICMS workshop on the Island.)

The first paper in the series is more than a foundational paper since (The) Fisher’s 1922 paper is about creating (almost) ex nihilo the field of (modern) mathematical statistics. I don’t know if there is any equivalence in other scientific disciplines of such an impact (and of such a man)… Roderick Little manages to convincingly engage with Fisher’s dismissive views on (not yet called) Bayesian analysis, although, to the latter’s defence, the formalisation of Bayesian inference at that time had not yet emerged. The second chapter is discussing Yates’ 1984 paper on tests for 2×2 contingency tables that he wrote 50 years after writing the original one in the first volume of JRSS. Roderick Little adds a detailed Bayesian analysis with the three standard reference priors, Jeffreys’ version proving quite close to Fisher’s exact test (conditional on both margins). The third chapter is aiming at the generic challenge of hypothesis testing, from the well-known opposition between Fisher and Neyman (both on the cover), to questioning the sanity of hard-set thresholds (with a mention of our American Statistician call to abandon (shi)p!). The later (thus) refers to the recent literature on the replicability crisis and the now famous ASA statement on p-values by Ron Wasserstein and Nicole Lazar, analysed in the chapter. But I would have like to read another full section on alternatives to hypothesis testing. While now a niche interest (imho), Fisher’s attempt at creating a posterior distribution without a prior, aka fiducial inference, is discussed in Chapter 4 with the Behrens-Fisher problem as the illustrating example. The chapter feels rather anticlimactic, with the comparison relying on the (Malay) Ghosh and Kim (2001) simulation results.

Birnbaum’s (1962) likelihood principle is the topic of Chapter 5 (and I cannot remember any of my students choosing this paper over the years, although there was at least one). Roderick Little recalls some sentences from the JASA discussion as an appetiser, a reminder of the time when these discussions could turn in scathing attacks. The chapter contains excerpts from Berger and Wolpert (1988)—which they were writing while I was spending a year at Purdue and which I have always recommended to my PhD students, albeit not for the classic seminar. It then moves to the controversies that surround this principle since its inception, in particular those accumulated by Deborah Mayo (also on the cover) as reported on the ‘Og. In the recent years, I have become less excited about the LP, in part due to the imprecision in its statement, which opens the door to conflicting interpretations. And in part due to the scarcity of models with non-trivial sufficient statistics. (I am also wondering if the sufficiency issue we highlighted in our ABC model choice criticism does relate to the mixture example at the end of the chapter.)

The next chapter is one all for compromise, through the calibrated Bayes perspective that credible statements should be close to confidence statements in the long run. Which I remember him presenting at ASC 2012. The concept is found in the very 1984 paper by Don Rubin (also on the cover) that contains the concept behind Approximate Bayesian Computation (ABC). And the chapter proceeds by listing strengths and weaknesses of frequentist and Bayesian perspectives, towards a fusion of both., e.g. though posterior predictive checks.

While the choice of a (general public) paper from Scientific American may sound surprising in Chapter 7, with Efron’s (on the cover) and Morris’ 1977 Stein’s paradox, I cannot but applaud, the more because this was the first paper I read when starting my PhD on the James-Stein estimators. Although this may sound like happening eons ago, the James and Stein (1961) paper—which is my age!—”created a considerable backlash” by toppling unbiasedness from its pedestal and exhibiting a paradox that 1+1+1≠3… Which Little reinterprets via a random effect (or Bayesian hierarchical) model. (And a chapter where I learned that Little’s father was a journalist, a characteristic he shared with Bruce Lindsay, as I found at Blonde, Glasgow, during an ICMS workshop). Relatedly, the next chapter is about the “57 varieties [of regression] paper” by Demptster, Schatzoff and Wermuth (1977). Apparently connected with Heinz 57 varieties of pickles. The paper considers Stein and ridge and variable selections versions for variable selection. The chapter also covers (Bayesian) Lasso and BART, as well as a brief all too brief mention of Spike & Slab priors—with my friend Veronika Ročková missing from the authors’ index!—,  but I was expecting from the title other, robust, forms of regression like L¹ regression and econometrics digressions. Chapter 10 can however been seen as a proxy since covering generalized estimating equations from a 1986 Biometrika paper of Liang and Zeger, with no Bayesian aspect (and an expected appearance of Communications in Statistics B).

Chapter 9 covers the almost immediately classic 1995 paper of Benjamini and Hochbeg on multiple regressions (that Series B turned into a discussion paper ten years later!). Although it spends more time on Berry’s (2012) recommendations than on FDR. The computational Chapter 11 brings together Efron’s (1979) bootstrap [with his picture on the cover] and MCMC, represented by the founding paper of Gelfand and Smith (1990, if mistakenly set in 1988 on p140). A bit of a strange mix imho as the former is more inferential than computational. And not giving the EM algorithm that much space. And not questioning MCMC methods as a good proxy to posterior distributions. Tukey’s Future of Data Analysis (as founding exploratory data analysis) and Breiman’s Two cultures (as launching statistical machine learning) meet in Chapter 12. (With a reminder that the latter invokes Occam’s razor—which may not be that appropriate for hugely overparameterised machine learning black boxes—and…the Rashomon principle! Meaning that distinct models may all fit the same data. Let me nitpickingly add the reference to Ryûnosuke Akutagawa as the author of Rashômon and other stories that Kurosawa adapted in his splendid movie). The chapter contains critical remarks from David Cox, Brad Efron, David Bickel, and Andrew Gelman, with a further section on Little’s view on modelling.

The last three chapters are on design and sampling, in connection with Little’s (and Rubin’s) works in the area. With a 1934 paper of Neyman (whose picture on the cover could have been chosen differently, albeit no fault of Neyman [or of Little!] that his toothbrush style of moustache dramatically got out of fashion!). With a return to calibrated Bayes and a reminiscence of Little’s time at the World Fertility Survey but (apparently) no mention of the probabilistic aspects of modern censuses (that saw my friends Steve Fienberg on the one side and Larry Brown and Marty Wells on the other side argue for and against it!), again relating to the reliance on statistical models. Chapter 14 relates randomized clinical trials to causality, which makes a (worthy) appearance there. Roderick Little also makes a clear case there against the retracted study linking vaccines and autism, a call that will unlikely not reach the current Trump administration and its Secretary of Health.

The book concludes with a list of twenty style and grammar suggestions for improved writing.

As should be crystal-clear from the above, I quite enjoyed the book and would definitely use its reading list in a graduate course whenever the opportunity arises. Once again, some choices are more personal to the author than others, and I would have place more emphasis on the fantastic Dawid, Stone and Zidek (1973)—with Jim Zidek also missing from the author index—, but all make sense in a walk through statistical classics. Let me however regret the absence therein of major actors like, e.g., D. Blackwell, C.R. Rao,  or G. Wahba (except in a stylistic example p199), two of whom were awarded the International Prize in Statistics.

[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE.]

ISBA 2016 [#5]

Posted in Mountains, pictures, Running, Statistics, Travel with tags , , , , , , , , , , , , , on June 18, 2016 by xi'an

from above Forte Village, Santa Magherita di Pula, Sardinia, June 17, 2016On Thursday, I started the day by a rather masochist run to the nearby hills, not only because of the very hour but also because, by following rabbit trails that were not intended for my size, I ended up being scratched by thorns and bramble all over!, but also with neat views of the coast around Pula.  From there, it was all downhill [joke]. The first morning talk I attended was by Paul Fearnhead and about efficient change point estimation (which is an NP hard problem or close to). The method relies on dynamic programming [which reminded me of one of my earliest Pascal codes about optimising a dam debit]. From my spectator’s perspective, I wonder[ed] at easier models, from Lasso optimisation to spline modelling followed by testing equality between bits. Later that morning, James Scott delivered the first Bayarri Lecture, created in honour of our friend Susie who passed away between the previous ISBA meeting and this one. James gave an impressive coverage of regularisation through three complex models, with the [hopefully not degraded by my translation] message that we should [as Bayesians] focus on important parts of those models and use non-Bayesian tools like regularisation. I can understand the practical constraints for doing so, but optimisation leads us away from a Bayesian handling of inference problems, by removing the ascertainment of uncertainty…

Later in the afternoon, I took part in the Bayesian foundations session, discussing the shortcomings of the Bayes factor and suggesting the use of mixtures instead. With rebuttals from [friends in] the audience!

This session also included a talk by Victor Peña and Jim Berger analysing and answering the recent criticisms of the Likelihood principle. I am not sure this answer will convince the critics, but I won’t comment further as I now see the debate as resulting from a vague notion of inference in Birnbaum‘s expression of the principle. Jan Hannig gave another foundation talk introducing fiducial distributions (a.k.a., Fisher’s Bayesian mimicry) but failing to provide a foundational argument for replacing Bayesian modelling. (Obviously, I am definitely prejudiced in this regard.)

The last session of the day was sponsored by BayesComp and saw talks by Natesh Pillai, Pierre Jacob, and Eric Xing. Natesh talked about his paper on accelerated MCMC recently published in JASA. Which surprisingly did not get discussed here, but would definitely deserve to be! As hopefully corrected within a few days, when I recoved from conference burnout!!! Pierre Jacob presented a work we are currently completing with Chris Holmes and Lawrence Murray on modularisation, inspired from the cut problem (as exposed by Plummer at MCMski IV in Chamonix). And Eric Xing spoke about embarrassingly parallel solutions, discussed a while ago here.

about the strong likelihood principle

Posted in Books, Statistics, University life with tags , , , , , , , on November 13, 2014 by xi'an

Deborah Mayo arXived a Statistical Science paper a few days ago, along with discussions by Jan Bjørnstad, Phil Dawid, Don Fraser, Michael Evans, Jan Hanning, R. Martin and C. Liu. I am very glad that this discussion paper came out and that it came out in Statistical Science, although I am rather surprised to find no discussion by Jim Berger or Robert Wolpert, and even though I still cannot entirely follow the deductive argument in the rejection of Birnbaum’s proof, just as in the earlier version in Error & Inference.  But I somehow do not feel like going again into a new debate about this critique of Birnbaum’s derivation. (Even though statements like the fact that the SLP “would preclude the use of sampling distributions” (p.227) would call for contradiction.)

“It is the imprecision in Birnbaum’s formulation that leads to a faulty impression of exactly what  is proved.” M. Evans

Indeed, at this stage, I fear that [for me] a more relevant issue is whether or not the debate does matter… At a logical cum foundational [and maybe cum historical] level, it makes perfect sense to uncover if and which if any of the myriad of Birnbaum’s likelihood Principles holds. [Although trying to uncover Birnbaum’s motives and positions over time may not be so relevant.] I think the paper and the discussions acknowledge that some version of the weak conditionality Principle does not imply some version of the strong likelihood Principle. With other logical implications remaining true. At a methodological level, I am less much less sure it matters. Each time I taught this notion, I got blank stares and incomprehension from my students, to the point I have now stopped altogether teaching the likelihood Principle in class. And most of my co-authors do not seem to care very much about it. At a purely mathematical level, I wonder if there even is ground for a debate since the notions involved can be defined in various imprecise ways, as pointed out by Michael Evans above and in his discussion. At a statistical level, sufficiency eventually is a strange notion in that it seems to make plenty of sense until one realises there is no interesting sufficiency outside exponential families. Just as there are very few parameter transforms for which unbiased estimators can be found. So I also spend very little time teaching and even less worrying about sufficiency. (As it happens, I taught the notion this morning!) At another and presumably more significant statistical level, what matters is information, e.g., conditioning means adding information (i.e., about which experiment has been used). While complex settings may prohibit the use of the entire information provided by the data, at a formal level there is no argument for not using the entire information, i.e. conditioning upon the entire data. (At a computational level, this is no longer true, witness ABC and similar limited information techniques. By the way, ABC demonstrates if needed why sampling distributions matter so much to Bayesian analysis.)

“Non-subjective Bayesians who (…) have to live with some violations of the likelihood principle (…) since their prior probability distributions are influenced by the sampling distribution.” D. Mayo (p.229)

In the end, the fact that the prior may depend on the form of the sampling distribution and hence does violate the likelihood Principle does not worry me so much. In most models I consider, the parameters are endogenous to those sampling distributions and do not live an ethereal existence independently from the model: they are substantiated and calibrated by the model itself, which makes the discussion about the LP rather vacuous. See, e.g., the coefficients of a linear model. In complex models, or in large datasets, it is even impossible to handle the whole data or the whole model and proxies have to be used instead, making worries about the structure of the (original) likelihood vacuous. I think we have now reached a stage of statistical inference where models are no longer accepted as ideal truth and where approximation is the hard reality, imposed by the massive amounts of data relentlessly calling for immediate processing. Hence, where the self-validation or invalidation of such approximations in terms of predictive performances is the relevant issue. Provided we can at all face the challenge…

who’s afraid of the big B wolf?

Posted in Books, Statistics, University life with tags , , , , , , , , , , on March 13, 2013 by xi'an

Aris Spanos just published a paper entitled “Who should be afraid of the Jeffreys-Lindley paradox?” in the journal Philosophy of Science. This piece is a continuation of the debate about frequentist versus llikelihoodist versus Bayesian (should it be Bayesianist?! or Laplacist?!) testing approaches, exposed in Mayo and Spanos’ Error and Inference, and discussed in several posts of the ‘Og. I started reading the paper in conjunction with a paper I am currently writing for a special volume in  honour of Dennis Lindley, paper that I will discuss later on the ‘Og…

“…the postdata severity evaluation (…) addresses the key problem with Fisherian p-values in the sense that the severity evaluation provides the “magnitude” of the warranted discrepancy from the null by taking into account the generic capacity of the test (that includes n) in question as it relates to the observed data”(p.88)

First, the antagonistic style of the paper is reminding me of Spanos’ previous works in that it relies on repeated value judgements (such as “Bayesian charge”, “blatant misinterpretation”, “Bayesian allegations that have undermined the credibility of frequentist statistics”, “both approaches are far from immune to fallacious interpretations”, “only crude rules of thumbs”, &tc.) and rhetorical sleights of hand. (See, e.g., “In contrast, the severity account ensures learning from data by employing trustworthy evidence (…), the reliability of evidence being calibrated in terms of the relevant error probabilities” [my stress].) Connectedly, Spanos often resorts to an unusual [at least for statisticians] vocabulary that amounts to newspeak. Here are some illustrations: “summoning the generic capacity of the test”, ‘substantively significant”, “custom tailoring the generic capacity of the test”, “the fallacy of acceptance”, “the relevance of the generic capacity of the particular test”, yes the term “generic capacity” is occurring there with a truly high frequency. Continue reading →

reading classics (#7)

Posted in Statistics with tags , , , , , , , , , on January 28, 2013 by xi'an

Last Monday, my student Li Chenlu presented the foundational 1962 JASA paper by Allan Birnbaum, On the Foundations of Statistical Inference. The very paper that derives the Likelihood Principle from the cumulated Conditional and Sufficiency principles and that had been discussed [maybe ad nauseam] on this ‘Og!!! Alas, thrice alas!, I was still stuck in the plane flying back from Atlanta as she was presenting her understanding of the paper, as the flight had been delayed four hours thanks to (or rather woe to!) the weather conditions in Paris the day before (chain reaction…):

I am sorry I could not attend this lecture and this for many reasons: first and  foremost, I wanted to attend every talk from my students both out of respect for them and to draw a comparison between their performances. My PhD student Sofia ran the seminar that day in my stead, for which I am quite grateful, but I do do wish I had been there… Second, this a.s. has been the most philosophical paper in the series.and I would have appreciated giving the proper light on the reasons for and the consequences of this paper as Li Chenlu stuck very much on the paper itself. (She provided additional references in the conclusion but they did not seem to impact the slides.)  Discussing for instance Berger’s and Wolpert’s (1988) new lights on the topic, as well as Deborah Mayo‘s (2010) attacks, and even Chang‘s (2012) misunderstandings, would have clearly helped the students.