Archive for Bayesian predictive

ISBA⁵

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , on July 4, 2026 by xi'an

After a pleasant run under the pouring rain (noisy rain that had not helped with my sleep or lack thereof), I attended both morning episodes of Calibrated Bayes, with a range of interesting, mostly novel, questions and solutions. And rekindling my interrogations about overparameterised models and model misspecification within a Bayesian framework. And the elusive notion of outliers.

The noon discussion around our report as the committee on the future of ISBA conferences went on rather well, given the circumstances, with about 50 participants with ideas and opinions on the opportunity of splitting the ISBA World into multiple hubs. And on the practical difficulties. (Despite the itch to do it, I did not intervene with my counter-objections to most of the raised objections! Nor mentioned BayesComp in Aussois.) Participants of the (unofficial) 2021 mirror in Marseille brought some welcomed support!

Daniele Durante’s Bayarri lecture on skew symmetric approximations was also attuned to the approximation spirit of the morning. The symmetrisation of the target reminded me of our (unpublished) folded method. With the (unrelated) question of the statistical meaning of a higher order Laplace expansion. And the calibration of this skewed approximation.

The final session on Cut Bayes couldn’t be missed!, since this has been an interest of mine’s since MCMSki IV in Chamonix! With an automated cut model construction by Robert Goude and an optimised choice of generalised Bayes posteriors by David Nott.

ISBA³

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , on July 2, 2026 by xi'an

After another absurdly early run along the Sendai River, in higher humidity than yesterday, I had plenty of time to have breakfast and commute to the conference centre for our privacy section. (“Ours” since I organised and chaired the session, but also all speakers were related with our ERC OCEAN project.) With talks by Antoine Luciano (Dauphine), Shenggang Hu (Warwick) and Joshua Bon (formerly Dauphine), Shenggang presenting his work with Gareth on calibrating a DP, noisy, unbiased, version of Metropolis-Hastings. Given the other sessions highly competing with ours, this 100% OCEAN session was well-attended! My second morning session was the theory & method part of the Savage Award, with unfortunately two of the speakers stuck in the US for visa issues and presenting remotely. There have been years where this session suffered from competition from parallel sessions, but this time showed a quite decent attendance.

As most members of the scientific committee of the BayesComp mirror in Aussois were present, we went out for lunch at a nearby hitsumabushi restaurant to plan registrations and local sessions. Which proved very productive despite enjoying the fantastic eel dishes! If making me miss the beginning of the afternoon session… and then the whole session as the room on predictive Bayes was packed (and some speakers had already delivered on the Isle of Skye). I managed to get to Bottond Szabo’s Foundation lecture, with again common threads with Skye, and then returned to my rental as I was quickly crashing…

approximately Bayes [on Skye]

Posted in Mountains, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , on May 28, 2026 by xi'an

Wow, what an exciting workshop in an equally exciting place! Strong themes were post-Bayes (Gibbs priors, martingale priors, predictive Bayes, &tc.) and deep neural network modelling. With animated discussions allowed by the free windows planned in the program. And the very early dinner at Sabhal Mòr Ostaig that let a long sunlit evening for impromptu Q&A’s [with a serving of lamb and another of haggis pie!]. Making me realise the large corpus of work I had missed in the past years on these topics, even though the satellite of BayesComp last year was already an eye opener. (Stay tuned for news about BayesComp 2027 & its mirror in Aussois!) The proposal in Jeff Miller’s discussion of Jeremias Knoblauch’s overview of post-Bayes [I’d rather favour another name!]  to consider directly likelihood values as the data was particularly appealing to me, while reminding me of the foundations of nested sampling. (Hopefully, a new perspective on uncertainty assessment for nested sampling is soon to be completed!)

On the non-academic side, the long days in The North helped with my running with above 90km bagged in the week (and no downpour on the runs). But little to my swimming since the water was cold enough to limit my laps to 5mn each time! Paradoxically the worst day was the one I chose for climbing the Inaccessible Pinnacle (as expanded in another ‘Og entry).

model uncertainty and missing data: an objective BAyesian perspective

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on September 16, 2025 by xi'an

My Spanish and objective Bayesian friends Gonzalo García-Donato, María Eugenia Castellanos, Stefano Cabras, Alicia Quirós, and Anabel Forte wrote an fairly exciting paper in BA that is open to discussion (for a few more days), to be discussed on 05 November (4:00 PM UTC | 11:00 AM EST | 5:00 PM CET).

The interplay between missing data and model uncertainty—two classic statistical problems—leads to primary questions that we formally address from an objective Bayesian perspective. For the general regression problem, we discuss the probabilistic justification of Rubin’s rules applied to the usual components of Bayesian variable selection, arguing that prior predictive marginals should be central to the pursued methodology. In the regression settings, we explore the conditions of prior distributions that make the missing data mechanism ignorable, provided that it is missing at random or completely at random. Moreover, when comparing multiple linear models, we provide a complete methodology for dealing with special cases, such as variable selection or uncertainty regarding model errors. In numerous simulation experiments, we demonstrate that our method outperforms or equals others, in consistently producing results close to those obtained using the full dataset. In general, the difference increases with the percentage of missing data and the correlation between the variables used for imputation.

The so-called Rubin’s identity is simply the representation of the posterior probability of a model γ given the observed data x⁰, p(γ|x⁰), as the integrated posterior probability of a model given both observed and latent data,  p(γ|x⁰, x¹), against the marginal of latent x¹ given observed x⁰. Since this marginal involves the probabilities p(γ|x⁰), this representation is not directly useful for a numerical implementation.

In this paper, missingness relates to some entries of either the covariates or the response variate. Which is less common but more realistic, especially if some covariates do not contribute to the response. (The missingness mechanism does not matter if the data is missing at random (à la Rubin). The computational solution (p9) is rather standard, simulating the missing variables given the observed variables. In my opinion, the elephant in the room is the super-delicate selection of a prior distribution on the missing covariates, as methinks this impacts in a considerable manner the actual value of the Bayes factor, hence the selection of the surviving model. (As a side remark, we are credited in Celeux et al. (2006) to have “extended DIC for missing data models or when missing data were present”, but our point was instead to point out the arbitrariness of the very definition of DIC in such contexts.)

“The standard Bayesian method for addressing the absence of prior information uses improper distributions. In estimation problems (the model is fixed), the impropriety of priors does not imply any additional difficulty as long as the posterior is proper” (p9)

The authors point out the well-known difficulty with improper priors but still resort to improper priors on the parameters shared by all models—which I dispute as being adequate, despite the arguments put forward on p15, right Haar measure or not—, while sticking to proper priors on the model-dependent parameters. Which unsurprisingly become Zellner’s g-priors. Or rather g’-priors, although the discussion seems to resolve into the (model-free) factor g’ being equal to 1 as for the g-priors. Again a strong term in the derivation of the Bayes factor.