Archive for Bayesian foundations

la vie (bayésienne), mode d’emploi

Posted in Books, pictures, Travel, University life with tags , , , , , , , , , , , on August 1, 2026 by xi'an

objective Bayesian inference now in paperback

Posted in Books, Statistics with tags , , , , , , , , , , , , , , on May 1, 2026 by xi'an

deep Bayes factor

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , on August 8, 2024 by xi'an

A recently arXived paper proposes an alternative approach to computing Bayes factors via deep learning, Deep Bayes Factors written by Jungeum Kim (presenting her work at JSM this very morning) and Veronika Ročková (whom I have known from her PhD years and whose COPSS Award we very gladly celebrated yesterday!). Which is obviously of interest to me, given my repeated visits to the challenge.

“we introduce Deep Bayes Factor (DeepBF), a neural classifier trained on simulated datasets to learn a mapping whose functional constitutes a Bayes factor estimator.”

Their approach is directly connected with various classification approaches to ABF, incl. the mythical inverse logistic version of Geyer (1994) and noise contrastive estimation of Gutmann and Hyvärinen (2010) (as well as our forested version). Which is called the likelihood-ratio trick here.

“Viewing the Bayes factor through the lens of binary classification aligns with Pudlo et al. (2016), who recast ABC model selection as a classification problem. They employ random forests to select a model by a majority vote. Instead, we focus on binary classification where the purpose is to learn marginal likelihood ratios.  Contrary to the method in Pudlo et al. (2016), our strategy circumvents a secondary learning phase for gauging model posterior estimates, delivering results in only one stage.”

The authors‘ solution stands with learning a classifier from simulated data from both models (and a basic log ratio utility), along iterations updating D from the gradient of the utility, the associated Bayes factor being the ratio D/(1-D) derived from the estimated classifier. There is a cost in producing new samples from the (same) predictives at each iteration (and I wonder if some recycling would be helpful, as well as reducing the sample size for the simpler model). In one of the remarks, the authors point out that “in the effort to see the best ABC performance, we intentionally use the full data Y as a summary statistic”, a remark that I find surprising given the overall consensus that the Bayes factor itself [when based on the full data] is close to optimal.

The method is overall consistent (in the data size n) under classical Bayesian asymptotics, sometimes even when the Bayes factor estimator is inconsistent, naturally expands to pseudo Bayes factors like intrinsic and fractional Bayes factors, also mileage varies in terms of numerical stability.

In the Bayesian model criticism section, the notion of opposing the actual dataset to a simulated one relates very much to Geyer’s (1994) solution. As well as to GANs, as noted in the paper. I did not look closely at the numerical comparisons in the experimental section, but they sound rich enough.

Masterclass in Bayesian Asymptotics, Université Paris Dauphine, 18-22 March 2024

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , on December 8, 2023 by xi'an

On the week of 18-22 March 2024, Judith Rousseau (Paris Dauphine & Oxford) will teach a Masterclass on Bayesian asymptotics. The masterclass takes place in Paris (on the PariSanté Campus) and consists of morning lectures and afternoon labs. Attendance is free with compulsory registration before 11 March (since the building is not accessible without prior registration).

The plan of the course is as follows

Part I: Parametric models
In this part, well- and mis-specified models will be considered.
– Asymptotic posterior distribution: asymptotic normality of the posterior,  penalization induced by the prior and the Bernstein von – Mises theorem. Regular and nonregular models will be treated.
– marginal likelihood and consistency of Bayes factors/model selection approaches.
– Empirical Bayes methods: asymptotic posterior distribution for parametric empirical Bayes methods.

Part II: Nonparametric and semiparametric models
– Posterior consistency and posterior convergence rates: statistical loss functions using the theory initiated by L. Schwartz and developed by Ghosal and Van der Vaart, results on less standard or well behaved losses.
– semiparametric Bernstein von Mises theorems.
– nonparametric Bernstein von Mises theorems and Uncertainty quantification.
– Stepping away from pure Bayes approaches: generalized Bayes, one step posteriors and cut posteriors.

Bayes’s theorem for improper mixtures

Posted in Books, Statistics, University life with tags , , , , , , on July 19, 2023 by xi'an

While looking for references for a Master summer project at Warwick on Bayesian inference on the Cauchy location parameter, I came across a 2011 Annals of Statistics paper by Peter McCullagh and Han Han.  Which expands the Bayesian framework to the improper case by considering a Poisson process over the parameter set with mean measure ν the improper prior. Instead of a single random parameter, this construct returns a countable collection of pairs (θ,y), while the observations induce a subset of that collection constrained by y∈A, a “sampling region” both capital to the derivation of the joint distribution and obscure in that A remains unspecified (but such that 0<ν(A)<∞ and conveniently returning the observed sample of y’s).

“Provided that the key finiteness condition is satisfied, this probabilistic analysis of the extended model may be interpreted as a vindication of improper Bayes procedures derived from the original model.”

“Thus, the existence of a joint probability model associated with an improper prior does not imply optimality in the form of coherence, consistency or admissibility.”

This is definitely fascinating!, even though I have troubles linking this infinite sequence of θ‘s with regular Bayesian inference, since the examples in the paper seem to revert to a single parameter value, as in §4.1, for the Normal model and §5 for the Cauchy model. The authors also revisit the marginalisation paradoxes of Dawid, Stone and Zidek (1973), with the argument that the improper measure leading to the paradox is not compatible with ν(A)<∞, hence does not define a natural conditional, while the “other” improper measure avoids the paradox.