Archive for Dickey-Savage ratio

e-values in Chennai

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , on July 23, 2025 by xi'an

To recap, I thus attended the BIRS-CMI workshop 25w5482 at the Chennai Mathematical Institute, Navalur, Tamil Nadu, in early July, for being intrigued by the developments around the concept. And enjoyed the week, from partaking in the company of friendly and enthusiastic academics to the exposure of new views and concepts, mostly remote from mine’s. Recall that an e-value attached to an hypothesis H described as a collection of distributions is a non-negative random variable E with expectation less than 1 for E~Q and all Q ∈ H. When a stopping rule is involved, the e-value is extended into an e-process. (Beyond Aaditya Ramdas’ E-book, Ruodu Wang also wrote a “tiny” review.) Aaditya Ramdas recalled in his introduction of the workshop that e-values are fundamentally equivalent to p-values and confidence intervals. And that a confidence sequence is a sequence of confidence intervals that contains the true value for all time steps t’s with a probability of at least 1-α.

The talks reflected a general belief in α levels and in Neyman-Pearsonian likelihood ratio optimality in simple vs simple settings, considering extension for sequential analysis settings, anytime inference, universality under general alternatives, and connections with FDRs, incl. Benjamini & Hochberg solution, but pointed out a lack of middle ground between frequentists and Bayesians.

“e-values have a clear interpretation in terms of betting and are closely related to likelihood ratios and other Bayes factor. At the same time, e–values do not require prior distributions conditional on the null and alternative hypotheses”

Although David R. Bickel attempted a Bayesian version, using a marginal likelihood ratio within betting settings, that is an incoming American Statistician paper. I may have being missing some aspects due to a lack of sleep the night before (!), but I find the attempt resulting in a fairly unusual vision of Bayesian testing as either not depending on any parameter or on the opposite using a family of priors. I did not understand either the “criticism” that the predictive depends on the prior and felt that this representation was bending in a rather onsiderable way the Bayesian perspective towards achieving a certain degree of agreement with p– and e-value notions, to conclude that the Bayes factor is an e-value. (As an aside, this may be the first paper that cited our critical review of Aitkin! Similarly, Shubhada Agrawal mentioned Roger Farrell in his talk, with whom we wrote a complete class Annals paper in the late 1980’s.) Nikos Ignatiadis also explored Empirical Bayes e-values, while Ben Chugg gave a presentation (constrained) admissibility, albeit under type-I error constraints that makes Bayes infeasible and using Neyman-Pearsonian loss functions. On the last day, Peter Grünwald tried for some BFF cohesion with openings on e-posteriors, treating hypothesis testing losses symmetrically, defining it as an inverse of e-values but incorporating pseudo-posteriors of many flavours like confidence, inferential, and fiducial distributions. He also mentioned a Savage-Dickey version while using an arbitrary prior, which is also an e-value, but with upper & lower meanings, again with measure issues

Given the hosting of the workshop in the Chennai Mathematical Institute, which is quite far from the centre of town (much closer to Mahabalipuram!), I did not visit Chennai but enjoyed the South Indian cuisine (albeit missing some fierceness in the spices!) and local fruits from street stands, if being sorry I could not find cocoa pods from nearby Kerala.

deep Bayes factor

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , on August 8, 2024 by xi'an

A recently arXived paper proposes an alternative approach to computing Bayes factors via deep learning, Deep Bayes Factors written by Jungeum Kim (presenting her work at JSM this very morning) and Veronika Ročková (whom I have known from her PhD years and whose COPSS Award we very gladly celebrated yesterday!). Which is obviously of interest to me, given my repeated visits to the challenge.

“we introduce Deep Bayes Factor (DeepBF), a neural classifier trained on simulated datasets to learn a mapping whose functional constitutes a Bayes factor estimator.”

Their approach is directly connected with various classification approaches to ABF, incl. the mythical inverse logistic version of Geyer (1994) and noise contrastive estimation of Gutmann and Hyvärinen (2010) (as well as our forested version). Which is called the likelihood-ratio trick here.

“Viewing the Bayes factor through the lens of binary classification aligns with Pudlo et al. (2016), who recast ABC model selection as a classification problem. They employ random forests to select a model by a majority vote. Instead, we focus on binary classification where the purpose is to learn marginal likelihood ratios.  Contrary to the method in Pudlo et al. (2016), our strategy circumvents a secondary learning phase for gauging model posterior estimates, delivering results in only one stage.”

The authors‘ solution stands with learning a classifier from simulated data from both models (and a basic log ratio utility), along iterations updating D from the gradient of the utility, the associated Bayes factor being the ratio D/(1-D) derived from the estimated classifier. There is a cost in producing new samples from the (same) predictives at each iteration (and I wonder if some recycling would be helpful, as well as reducing the sample size for the simpler model). In one of the remarks, the authors point out that “in the effort to see the best ABC performance, we intentionally use the full data Y as a summary statistic”, a remark that I find surprising given the overall consensus that the Bayes factor itself [when based on the full data] is close to optimal.

The method is overall consistent (in the data size n) under classical Bayesian asymptotics, sometimes even when the Bayes factor estimator is inconsistent, naturally expands to pseudo Bayes factors like intrinsic and fractional Bayes factors, also mileage varies in terms of numerical stability.

In the Bayesian model criticism section, the notion of opposing the actual dataset to a simulated one relates very much to Geyer’s (1994) solution. As well as to GANs, as noted in the paper. I did not look closely at the numerical comparisons in the experimental section, but they sound rich enough.

Bayes factors revisited

Posted in Books, Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , on March 22, 2021 by xi'an

 

“Bayes factor analyses are highly sensitive to and crucially depend on prior assumptions about model parameters (…) Note that the dependency of Bayes factors on the prior goes beyond the dependency of the posterior on the prior. Importantly, for most interesting problems and models, Bayes factors cannot be computed analytically.”

Daniel J. Schad, Bruno Nicenboim, Paul-Christian Bürkner, Michael Betancourt, Shravan Vasishth have just arXived a massive document on the Bayes factor, worrying about the computation of this common tool, but also at the variability of decisions based on Bayes factors, e.g., stressing correctly that

“…we should not confuse inferences with decisions. Bayes factors provide inference on hypotheses. However, to obtain discrete decisions (…) from continuous inferences in a principled way requires utility functions. Common decision heuristics (e.g., using Bayes factor larger than 10 as a discovery threshold) do not provide a principled way to perform decisions, but are merely heuristic conventions.”

The text is long and at times meandering (at least in the sections I read), while trying a wee bit too hard to bring up the advantages of using Bayes factors versus frequentist or likelihood solutions. (The likelihood ratio being presented as a “frequentist” solution, which I think is an incorrect characterisation.) For instance, the starting point of preferring a model with a higher marginal likelihood is presented as an evidence (oops!) rather than argumented. Since this quantity depends on both the prior and the likelihood, it being high or low is impacted by both. One could then argue that using its numerical value as an absolute criterion amounts to selecting the prior a posteriori as much as checking the fit to the data! The paper also resorts to the Occam’s razor argument, which I wish we could omit, as it is a vague criterion, wide open to misappropriation. It is also qualitative, rather than quantitative, hence does not help in calibrating the Bayes factor.

Concerning the actual computation of the Bayes factor, an issue that has always been a concern and a research topic for me, the authors consider only two “very common methods”, the Savage–Dickey density ratio method and bridge sampling. We discussed the shortcomings of the Savage–Dickey density ratio method with Jean-Michel Marin about ten years ago. And while bridge sampling is an efficient approach when comparing models of the same dimension, I have reservations about this efficiency in other settings. Alternative approaches like importance nested sampling, noise contrasting estimation or SMC samplers are often performing quite efficiently as normalising constant approximations. (Not to mention our version of harmonic mean estimator with HPD support.)

Simulation-based inference is based on the notion that simulated data can be produced from the predictive distributions. Reminding me of ABC model choice to some extent. But I am uncertain this approach can be used to calibrate the decision procedure to select the most appropriate model. We thought about using this approach in our testing by mixture paper and it is favouring the more complex of the two models. This seems also to occur for the example behind Figure 5 in the paper.

Two other points: first, the paper does not consider the important issue with improper priors, which are not rigorously compatible with Bayes factors, as I discussed often in the past. And second, Bayes factors are not truly Bayesian decision procedures, since they remove the prior weights on the models, thus the mention of utility functions therein seems inappropriate unless a genuine utility function can be produced.

O’Bayes 19/4

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , on July 4, 2019 by xi'an

Last talks of the conference! With Rui Paulo (along with Gonzalo Garcia-Donato) considering the special case of factors when doing variable selection. Which is an interesting question that I had never considered, as at best I would remove all leves or keeping them all. Except that there may be misspecification in the factors as for instance when several levels have the same impact.With Michael Evans discussing a paper that he wrote for the conference! Following his own approach to statistical evidence. And including his reluctance to cover infinity (calling on Gauß for backup!) or continuity, and his call to falsify a Bayesian model by checking it can be contradicted by the data. His assumption that checking for prior is separable from checking for [sampling] model is debatable. (With another mention made of the Savage-Dickey ratio.)

And with Dimitris Fouskakis giving a wide ranging assessment [which Mark Steel (Warwick) called a PEP talk!] of power-expected-posterior priors, used with reference (and usually improper) priors. Which in retrospect would have suited better the beginning of the conference as it provided a background to several of the talks. Raising a question (from my perspective) on using the maximum likelihood estimator as a pseudo-sufficient statistic when this MLE is computed for the base (simplest) model. Maybe an ABC induced bias in this question as it would not work for ABC model choice.

Overall, I think the scientific outcomes of the conference were quite positive: a wide range of topics and perspectives, a reasonable and diverse attendance, especially when considering the heavy load of related conferences in the surrounding weeks (the “June fatigue”!), animated poster sessions. I am obviously not the one to assess the organisation of the conference! Things I forgot to do in this regard: organise transportation from Oxford to Warwick University, provide an attached room for in-pair research, insist on sustainability despite the imposed catering solution, facilitate sharing joint transportation to and from the Warwick campus, mention that tap water was potable, and… wear long pants when running in nettles.

O’Bayes 19/3

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , on July 2, 2019 by xi'an

Nancy Reid gave the first talk of the [Canada] day, in an impressive comparison of all approaches in statistics that involve a distribution of sorts on the parameter, connected with the presentation she gave at BFF4 in Harvard two years ago, including safe Bayes options this time. This was related to several (most?) of the talks at the conference, given the level of worry (!) about the choice of a prior distribution. But the main assessment of the methods still seemed to be centred on a frequentist notion of calibration, meaning that epistemic interpretations of probabilities and hence most of Bayesian answers were disqualified from the start.

In connection with Nancy’s focus, Peter Hoff’s talk also concentrated on frequency valid confidence intervals in (linear) hierarchical models. Using prior information or structure to build better and shrinkage-like confidence intervals at a given confidence level. But not in the decision-theoretic way adopted by George Casella, Bill Strawderman and others in the 1980’s. And also making me wonder at the relevance of contemplating a fixed coverage as a natural goal. Above, a side result shown by Peter that I did not know and which may prove useful for Monte Carlo simulation.

Jaeyong Lee worked on a complex model for banded matrices that starts with a regular Wishart prior on the unrestricted space of matrices, computes the posterior and then projects this distribution onto the constrained subspace. (There is a rather consequent literature on this subject, including works by David Dunson in the past decade of which I was unaware.) This is a smart demarginalisation idea but I wonder a wee bit at the notion as the constrained space has measure zero for the larger model. This could explain for the resulting posterior not being a true posterior for the constrained model in the sense that there is no prior over the constrained space that could return such a posterior. Another form of marginalisation paradox. The crux of the paper is however about constructing a functional form of minimaxity. In his discussion of the paper, Guido Consonni provided a representation of the post-processed posterior (P³) that involves the Dickey-Savage ratio, sort of, making me more convinced of the connection.

As a lighter aside, one item of local information I should definitely have broadcasted more loudly and long enough in advance to the conference participants is that the University of Warwick is not located in ye olde town of Warwick, where there is no university, but on the outskirts of the city of Coventry, but not to be confused with the University of Coventry. Located in Coventry.