Archive for FDRs

e-values in Chennai

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , on July 23, 2025 by xi'an

To recap, I thus attended the BIRS-CMI workshop 25w5482 at the Chennai Mathematical Institute, Navalur, Tamil Nadu, in early July, for being intrigued by the developments around the concept. And enjoyed the week, from partaking in the company of friendly and enthusiastic academics to the exposure of new views and concepts, mostly remote from mine’s. Recall that an e-value attached to an hypothesis H described as a collection of distributions is a non-negative random variable E with expectation less than 1 for E~Q and all Q ∈ H. When a stopping rule is involved, the e-value is extended into an e-process. (Beyond Aaditya Ramdas’ E-book, Ruodu Wang also wrote a “tiny” review.) Aaditya Ramdas recalled in his introduction of the workshop that e-values are fundamentally equivalent to p-values and confidence intervals. And that a confidence sequence is a sequence of confidence intervals that contains the true value for all time steps t’s with a probability of at least 1-α.

The talks reflected a general belief in α levels and in Neyman-Pearsonian likelihood ratio optimality in simple vs simple settings, considering extension for sequential analysis settings, anytime inference, universality under general alternatives, and connections with FDRs, incl. Benjamini & Hochberg solution, but pointed out a lack of middle ground between frequentists and Bayesians.

“e-values have a clear interpretation in terms of betting and are closely related to likelihood ratios and other Bayes factor. At the same time, e–values do not require prior distributions conditional on the null and alternative hypotheses”

Although David R. Bickel attempted a Bayesian version, using a marginal likelihood ratio within betting settings, that is an incoming American Statistician paper. I may have being missing some aspects due to a lack of sleep the night before (!), but I find the attempt resulting in a fairly unusual vision of Bayesian testing as either not depending on any parameter or on the opposite using a family of priors. I did not understand either the “criticism” that the predictive depends on the prior and felt that this representation was bending in a rather onsiderable way the Bayesian perspective towards achieving a certain degree of agreement with p– and e-value notions, to conclude that the Bayes factor is an e-value. (As an aside, this may be the first paper that cited our critical review of Aitkin! Similarly, Shubhada Agrawal mentioned Roger Farrell in his talk, with whom we wrote a complete class Annals paper in the late 1980’s.) Nikos Ignatiadis also explored Empirical Bayes e-values, while Ben Chugg gave a presentation (constrained) admissibility, albeit under type-I error constraints that makes Bayes infeasible and using Neyman-Pearsonian loss functions. On the last day, Peter Grünwald tried for some BFF cohesion with openings on e-posteriors, treating hypothesis testing losses symmetrically, defining it as an inverse of e-values but incorporating pseudo-posteriors of many flavours like confidence, inferential, and fiducial distributions. He also mentioned a Savage-Dickey version while using an arbitrary prior, which is also an e-value, but with upper & lower meanings, again with measure issues

Given the hosting of the workshop in the Chennai Mathematical Institute, which is quite far from the centre of town (much closer to Mahabalipuram!), I did not visit Chennai but enjoyed the South Indian cuisine (albeit missing some fierceness in the spices!) and local fruits from street stands, if being sorry I could not find cocoa pods from nearby Kerala.

a Bayesian interpretation of FDRs?

Posted in Statistics with tags , , , , , , , , , , on April 12, 2018 by xi'an

This week, I happened to re-read John Storey’ 2003 “The positive discovery rate: a Bayesian interpretation and the q-value”, because I wanted to check a connection with our testing by mixture [still in limbo] paper. I however failed to find what I was looking for because I could not find any Bayesian flavour in the paper apart from an FRD expressed as a “posterior probability” of the null, in the sense that the setting was one of opposing two simple hypotheses. When there is an unknown parameter common to the multiple hypotheses being tested, a prior distribution on the parameter makes these multiple hypotheses connected. What makes the connection puzzling is the assumption that the observed statistics defining the significance region are independent (Theorem 1). And it seems to depend on the choice of the significance region, which should be induced by the Bayesian modelling, not the opposite. (This alternative explanation does not help either, maybe because it is on baseball… Or maybe because the sentence “If a player’s [posterior mean] is above .3, it’s more likely than not that their true average is as well” does not seem to appear naturally from a Bayesian formulation.) [Disclaimer: I am not hinting at anything wrong or objectionable in Storey’s paper, just being puzzled by the Bayesian tag!]

robust Bayesian FDR control with Bayes factors [a reply]

Posted in Statistics, University life with tags , , , , on January 17, 2014 by xi'an

(Following my earlier discussion of his paper, Xiaoquan Wen sent me this detailed reply.)

I think it is appropriate to start my response to your comments by introducing a little bit of the background information on my research interest and the project itself: I consider myself as an applied statistician, not a theorist, and I am interested in developing theoretically sound and computationally efficient methods to solve practical problems. The FDR project originated from a practical application in genomics involving hypothesis testing. The details of this particular application can be found in this published paper, and the simulations in the manuscript are also designed for a similar context. In this application, the null model is trivially defined, however there exist finitely many alternative scenarios for each test. We proposed a Bayesian solution that handles this complex setting quite nicely: in brief, we chose to model each possible alternative scenario parametrically, and by taking advantage of Bayesian model averaging, Bayes factor naturally ended up as our test statistic. We had no problem in demonstrating the resulting Bayes factor is much more powerful than the existing approaches, even accounting for the prior (mis-)modeling for Bayes factors. However, in this genomics application, there are potentially tens of thousands of tests need to be simultaneously performed, and FDR control becomes necessary and challenging. Continue reading →