Archive for empirical Bayes methods

e-values in Chennai

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , on July 23, 2025 by xi'an

To recap, I thus attended the BIRS-CMI workshop 25w5482 at the Chennai Mathematical Institute, Navalur, Tamil Nadu, in early July, for being intrigued by the developments around the concept. And enjoyed the week, from partaking in the company of friendly and enthusiastic academics to the exposure of new views and concepts, mostly remote from mine’s. Recall that an e-value attached to an hypothesis H described as a collection of distributions is a non-negative random variable E with expectation less than 1 for E~Q and all Q ∈ H. When a stopping rule is involved, the e-value is extended into an e-process. (Beyond Aaditya Ramdas’ E-book, Ruodu Wang also wrote a “tiny” review.) Aaditya Ramdas recalled in his introduction of the workshop that e-values are fundamentally equivalent to p-values and confidence intervals. And that a confidence sequence is a sequence of confidence intervals that contains the true value for all time steps t’s with a probability of at least 1-α.

The talks reflected a general belief in α levels and in Neyman-Pearsonian likelihood ratio optimality in simple vs simple settings, considering extension for sequential analysis settings, anytime inference, universality under general alternatives, and connections with FDRs, incl. Benjamini & Hochberg solution, but pointed out a lack of middle ground between frequentists and Bayesians.

“e-values have a clear interpretation in terms of betting and are closely related to likelihood ratios and other Bayes factor. At the same time, e–values do not require prior distributions conditional on the null and alternative hypotheses”

Although David R. Bickel attempted a Bayesian version, using a marginal likelihood ratio within betting settings, that is an incoming American Statistician paper. I may have being missing some aspects due to a lack of sleep the night before (!), but I find the attempt resulting in a fairly unusual vision of Bayesian testing as either not depending on any parameter or on the opposite using a family of priors. I did not understand either the “criticism” that the predictive depends on the prior and felt that this representation was bending in a rather onsiderable way the Bayesian perspective towards achieving a certain degree of agreement with p– and e-value notions, to conclude that the Bayes factor is an e-value. (As an aside, this may be the first paper that cited our critical review of Aitkin! Similarly, Shubhada Agrawal mentioned Roger Farrell in his talk, with whom we wrote a complete class Annals paper in the late 1980’s.) Nikos Ignatiadis also explored Empirical Bayes e-values, while Ben Chugg gave a presentation (constrained) admissibility, albeit under type-I error constraints that makes Bayes infeasible and using Neyman-Pearsonian loss functions. On the last day, Peter Grünwald tried for some BFF cohesion with openings on e-posteriors, treating hypothesis testing losses symmetrically, defining it as an inverse of e-values but incorporating pseudo-posteriors of many flavours like confidence, inferential, and fiducial distributions. He also mentioned a Savage-Dickey version while using an arbitrary prior, which is also an e-value, but with upper & lower meanings, again with measure issues

Given the hosting of the workshop in the Chennai Mathematical Institute, which is quite far from the centre of town (much closer to Mahabalipuram!), I did not visit Chennai but enjoyed the South Indian cuisine (albeit missing some fierceness in the spices!) and local fruits from street stands, if being sorry I could not find cocoa pods from nearby Kerala.

Masterclass in Bayesian Asymptotics, Université Paris Dauphine, 18-22 March 2024

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , on December 8, 2023 by xi'an

On the week of 18-22 March 2024, Judith Rousseau (Paris Dauphine & Oxford) will teach a Masterclass on Bayesian asymptotics. The masterclass takes place in Paris (on the PariSanté Campus) and consists of morning lectures and afternoon labs. Attendance is free with compulsory registration before 11 March (since the building is not accessible without prior registration).

The plan of the course is as follows

Part I: Parametric models
In this part, well- and mis-specified models will be considered.
– Asymptotic posterior distribution: asymptotic normality of the posterior,  penalization induced by the prior and the Bernstein von – Mises theorem. Regular and nonregular models will be treated.
– marginal likelihood and consistency of Bayes factors/model selection approaches.
– Empirical Bayes methods: asymptotic posterior distribution for parametric empirical Bayes methods.

Part II: Nonparametric and semiparametric models
– Posterior consistency and posterior convergence rates: statistical loss functions using the theory initiated by L. Schwartz and developed by Ghosal and Van der Vaart, results on less standard or well behaved losses.
– semiparametric Bernstein von Mises theorems.
– nonparametric Bernstein von Mises theorems and Uncertainty quantification.
– Stepping away from pure Bayes approaches: generalized Bayes, one step posteriors and cut posteriors.

Approximation Methods in Bayesian Analysis [#2]

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on June 22, 2023 by xi'an

A more theoretical Day #2 of the workshop, with Debdeep Pati comparing two representations of Gaussian processes with significantly different efficiencies, and Aad van der Vaart presenting a form of linearisation for a range of inverse problems, Kolyan Ray debiasing Lasso impacts by variational Bayes, although through a somewhat intricate process that distanced the procedure from Bayesian grounds imho, Judith Rousseau (Dauphine) also drifting away from Bayesian canons by looking anew at empirical Bayes with surprising differences from genuine B analysis, connecting with the cutoff phenomenon she and Kerrie exhibited in their 2011 mixture paper, as well as labelling the marginal likelihood a misspecified model. Trevor Campbell and Sinead Williamson both provided Bayesian perspectives on normalising flows, in particular the impact of computer imprecision on reversibility, leading to the notion of shadow paths (screenshot below), while Giovanni Rebaudo talked about mixtures supported by trees, a fascinating object!

On Day #3, Marc Beaumont talked on a mixture of composite likelihood à la Ryden, making me wonder of optimisation of blocks for HMC? EP-ABC, with the issue of the unknown amount of approximation, and adaptivity?, Maria de Iorio presented work on finite and infinite mixtures with repulsive (Coulomb) priors, achieving a unified framework, plus known evidence (?), with a correlated talk by Federico Camerlenghi in the afternoon, with novel notions (for me) of Palm measures and calculus, and another correlated talk by María-Fernanda Gil Leyva Villa, on stick-breaking processes for species sampling with dependent length variables, with related improvements in Gibbs implementation (screenshot below).
This was followed by two theoretical talks on continuous time processes by Paul Jenkins (Warwick) on the fine properties of the Flemming-Viot process, with mentions of Don Dawson’s results reminding me of the 1988 and 1989 summers I spent at Carleton University, where he was located at the time, and Matteo Ruggieri, with the novel (to me) notion of dual Markov processes that could prove useful in a lot of latent variable models. Fabrizio Leisen expanded on his early work on partial exchangeability and Steve MacEachern on dependent quantile pyramids, which relate to quantile regression, a constant source of puzzlement for me. Motivating the perspective by robustness and misspecification arguments. But I am a wee bit puzzled by the distinction between quantile pyramids and other non-parametric solutions.

On the outdoor front (in early mornings), choppy waters at sea (in the Sugiton calanque, pictured above) thanks to the endless mistral wind, nice run down from Mont Puget with friends, limited utility of my rented mountain bike (except to reach the nearest supermarket, 3km away)

Finite mixture models do not reliably learn the number of components

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , on October 15, 2022 by xi'an

When preparing my talk for Padova, I found that Diana Cai, Trevor Campbell, and Tamara Broderick wrote this ICML / PLMR paper last year on the impossible estimation of the number of components in a mixture.

“A natural check on a Bayesian mixture analysis is to establish that the Bayesian posterior on the number of components increasingly concentrates near the truth as the number of data points becomes arbitrarily large.” Cai, Campbell & Broderick (2021)

Which seems to contradict [my formerly-Glaswegian friend] Agostino Nobile  who showed in his thesis that the posterior on the number of components does concentrate at the true number of components, provided the prior contains that number in its support. As well as numerous papers on the consistency of the Bayes factor, including the one against an infinite mixture alternative, as we discussed in our recent paper with Adrien and Judith. And reminded me of the rebuke I got in 2001 from the late David McKay when mentioning that I did not believe in estimating the number of components, both because of the impact of the prior modelling and of the tendency of the data to push for more clusters as the sample size increased. (This was a most lively workshop Mike Titterington and I organised at ICMS in Edinburgh, where Radford Neal also delivered an impromptu talk to argue against using the Galaxy dataset as a benchmark!)

“In principle, the Bayes factor for the MFM versus the DPM could be used as an empirical criterion for choosing between the two models, and in fact, it is quite easy to compute an approximation to the Bayes factor using importance sampling” Miller & Harrison (2018)

This is however a point made in Miller & Harrison (2018) that the estimation of k logically goes south if the data is not from the assumed mixture model. In this paper, Cai et al. demonstrate that the posterior diverges, even when it depends on the sample size. Or even the sample as in empirical Bayes solutions.

21w5107 [½day 3]

Posted in pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , on December 2, 2021 by xi'an

Day [or half-day] three started without firecrackers and with David Rossell (formerly Warwick) presenting an empirical Bayes approach to generalised linear model choice with a high degree of confounding, using approximate Laplace approximations. With considerable improvements in the experimental RMSE. Making feeling sorry there was no apparent fully (and objective?) Bayesian alternative! (Two more papers on my reading list that I should have read way earlier!) Then Veronika Rockova discussed her work on approximate Metropolis-Hastings by classification. (With only a slight overlap with her One World ABC seminar.) Making me once more think of Geyer’s n⁰564 technical report, namely the estimation of a marginal likelihood by a logistic discrimination representation. Her ABC resolution replaces the tolerance step by an exponential of minus the estimated Kullback-Leibler divergence between the data density and the density associated with the current value of the parameter. (I wonder if there is a residual multiplicative constant there… Presumably not. Great idea!) The classification step need be run at every iteration, which could be sped up by subsampling.

On the always fascinating theme of loss based posteriors, à la Bissiri et al., Jack Jewson (formerly Warwick) exposed his work generalised Bayesian and improper models (from Birmingham!). Using data to decide between model and loss, which sounds highly unorthodox! First difficulty is that losses are unscaled. Or even not integrable after an exponential transform. Hence the notion of improper models. As in the case of robust Tukey’s loss, which is bounded by an arbitrary κ. Immediately I wonder if the fact that the pseudo-likelihood does not integrate is important beyond the (obvious) absence of a normalising constant. And the fact that this is not a generative model. And the answer came a few slides later with the use of the Hyvärinen score. Rather than the likelihood score. Which can itself be turned into a H-posterior, very cool indeed! Although I wonder at the feasibility of finding an [objective] prior on κ.

Rajesh Ranganath completed the morning session with a talk on [the difficulty of] connecting Bayesian models and complex prediction models. Using instead a game theoretic approach with Brier scores under censoring. While there was a connection with Veronika’s use of a discriminator as a likelihood approximation, I had trouble catching the overall message…