Archive for point null hypotheses

a lesser-known correlate of the Jeffreys-Lindley paradox (with discussion)

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , on October 19, 2024 by xi'an

Two UBC faculty, Harlan Campbell and Paul Gustafson, wrote a paper entitled “Defining a Credible Interval Is Not Always Possible with “Point-Null” Priors: A Lesser-Known Correlate of the Jeffreys-Lindley Paradox” in Bayesian Analysis (2024, 19, Number 3, pp. 925–984), which got discussed and presented on the BA webinar yesterday. I missed the call for discussion, on a topic I would have liked very much to discuss and an analysis I strongly disagree with. Fortunately, several of the discussants in the webinar and in the printed version advanced some of my points (as. e.g., Bertrand Clarke in the above slide screen-shot from the on-line video).

I find the paper somewhat missing in linking with the history of the topic, with no mention of Berger & Sellke (1987) that comes as a counterpoint to Casella &—the other—Berger (1987), opposing one sided to two sided tests. Or of matching priors, which connect credible and confidence intervals to higher orders. But the central issue with the apparent contradiction between rejecting the point null hypothesis and returning a credible interval that contains the null is that the construction proceeds from a model averaged posterior. Which fundamentally contradicts the construct of a pair of priors attached with each model towards selecting the fittest one. And requires a far-from-innocent choice of respective prior weights for both models, an ill-defined notion I have repeatedly criticised here and elsewhere. Model averaging clashes with model selection in both decision-theoretic and modelling terms. In model averaging terms, the disappearance of the opposition exhibited by the authors in the predictive distribution, as shown by discussants Held and Pawel, is unsurprising. And makes the spike-and-slab prior far of a necessity. Contrariwise to the model selection case where it proves unavoidable. And for which a merged credible interval does not make sense (to me at least) since it should be constructed once one (and only one) of the two models is chosen. At this point, that the other model ever was considered should not impact subsequent inference. And within that perspective I do not see the relevance of agnostic (ignoring the model choice ation) 5% confidence or credible regions.

“…considers the regime of a fixed true parameter value as n increases [and] of a fixed p-value…” (p928)

With regards with the connection with the Jeffreys-Lindley (or Lindley-Jeffreys) so-called paradox, on which I have already written a lot (or even too much!), many of the earlier objections resurface. Like the measure-theoretic difficulty in including within a continuous interval an atom, i.e., a value with a point mass. Which isolates this atom away from any other value in the interval (and of course creates discontinuities). Or fixing the p-value forever after (when n goes to infinity), as in the graph below (p929). Or treating an improper prior without further caution than with a proper prior. Especially when these are “created” by the decision problem itself.

 

demystify Lindley’s paradox [or not]

Posted in Statistics with tags , , , , , on March 18, 2020 by xi'an

Another paper on Lindley’s paradox appeared on arXiv yesterday, by Guosheng Yin and Haolun Shi, interpreting posterior probabilities as p-values. The core of this resolution is to express a two-sided hypothesis as a combination of two one-sided hypotheses along the opposite direction, taking then advantage of the near equivalence of posterior probabilities under some non-informative prior and p-values in the later case. As already noted by George Casella and Roger Berger (1987) and presumably earlier. The point is that one-sided hypotheses are quite friendly to improper priors, since they only require a single prior distribution. Rather than two when point nulls are under consideration. The p-value created by merging both one-sided hypotheses makes little sense to me as it means testing that both θ≥0 and θ≤0, resulting in the proposal of a p-value that is twice the minimum of the one-sided p-values, maybe due to a Bonferroni correction, although the true value should be zero… I thus see little support for this approach to resolving Lindley paradox in that it bypasses the toxic nature of point-null hypotheses that require a change of prior toward a mixture supporting one hypothesis and the other. Here the posterior of the point-null hypothesis is defined in exactly the same way the p-value is defined, hence making the outcome most favourable to the agreement but not truly addressing the issue.

unrejected null [xkcd]

Posted in Statistics with tags , , , , , on July 18, 2018 by xi'an

estimation versus testing [again!]

Posted in Books, Statistics, University life with tags , , , , , , , , , , on March 30, 2017 by xi'an

The following text is a review I wrote of the paper “Parameter estimation and Bayes factors”, written by J. Rouder, J. Haff, and J. Vandekerckhove. (As the journal to which it is submitted gave me the option to sign my review.)

The opposition between estimation and testing as a matter of prior modelling rather than inferential goals is quite unusual in the Bayesian literature. In particular, if one follows Bayesian decision theory as in Berger (1985) there is no such opposition, but rather the use of different loss functions for different inference purposes, while the Bayesian model remains single and unitarian.

Following Jeffreys (1939), it sounds more congenial to the Bayesian spirit to return the posterior probability of an hypothesis H⁰ as an answer to the question whether this hypothesis holds or does not hold. This however proves impossible when the “null” hypothesis H⁰ has prior mass equal to zero (or is not measurable under the prior). In such a case the mathematical answer is a probability of zero, which may not satisfy the experimenter who asked the question. More fundamentally, the said prior proves inadequate to answer the question and hence to incorporate the information contained in this very question. This is how Jeffreys (1939) justifies the move from the original (and deficient) prior to one that puts some weight on the null (hypothesis) space. It is often argued that the move is unnatural and that the null space does not make sense, but this only applies when believing very strongly in the model itself. When considering the issue from a modelling perspective, accepting the null H⁰ means using a new model to represent the model and hence testing becomes a model choice problem, namely whether or not one should use a complex or simplified model to represent the generation of the data. This is somehow the “unification” advanced in the current paper, albeit it does appear originally in Jeffreys (1939) [and then numerous others] rather than the relatively recent Mitchell & Beauchamp (1988). Who may have launched the spike & slab denomination.

I have trouble with the analogy drawn in the paper between the spike & slab estimate and the Stein effect. While the posterior mean derived from the spike & slab posterior is indeed a quantity drawn towards zero by the Dirac mass at zero, it is rarely the point in using a spike & slab prior, since this point estimate does not lead to a conclusion about the hypothesis: for one thing it is never exactly zero (if zero corresponds to the null). For another thing, the construction of the spike & slab prior is both artificial and dependent on the weights given to the spike and to the slab, respectively, to borrow expressions from the paper. This approach thus leads to model averaging rather than hypothesis testing or model choice and therefore fails to answer the (possibly absurd) question as to which model to choose. Or refuse to choose. But there are cases when a decision must be made, like continuing a clinical trial or putting a new product on the market. Or not.

In conclusion, the paper surprisingly bypasses the decision-making aspect of testing and hence ends up with a inconclusive setting, staying midstream between Bayes factors and credible intervals. And failing to provide a tool for decision making. The paper also fails to acknowledge the strong dependence of the Bayes factor on the tail behaviour of the prior(s), which cannot be [completely] corrected by a finite sample, hence its relativity and the unreasonableness of a fixed scale like Jeffreys’ (1939).

Measuring statistical evidence using relative belief [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , on July 22, 2015 by xi'an

“It is necessary to be vigilant to ensure that attempts to be mathematically general do not lead us to introduce absurdities into discussions of inference.” (p.8)

This new book by Michael Evans (Toronto) summarises his views on statistical evidence (expanded in a large number of papers), which are a quite unique mix of Bayesian  principles and less-Bayesian methodologies. I am quite glad I could receive a version of the book before it was published by CRC Press, thanks to Rob Carver (and Keith O’Rourke for warning me about it). [Warning: this is a rather long review and post, so readers may chose to opt out now!]

“The Bayes factor does not behave appropriately as a measure of belief, but it does behave appropriately as a measure of evidence.” (p.87)

Continue reading →