Archive for Bayesian nonparametrics

JSM 2024, Portland, miniday 4

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , on August 11, 2024 by xi'an

Final (half)day at JSM is always a sad thing as most people have left, people are busy dismantling booths and packing boxes, and the few remaining participants are fidgety and sitting on their suitcases (there may have been more suitcases than people in the main area that day!), cafés are minimally staffed or simply closed. Hence not the best time-slot to deliver one’s talk! Still, a few dozen people attended our session. The session topic was Bridge the Gap: Differential Privacy and Statistical Analysis, organised by Bei Jiang whom I met last summer in Kelowna, at a BIRS workshop on privacy. Where I spoke on setting up a complete decision-theoretic framework, as developed (and still in development) within our Ocean group, esp. Joshua Bon, Stan du Ché, and Judith Rousseau. (Rather than on the original plan of talking about convergence versus privacy, as a criticism of differential privacy.) The other talks were by Shurong Li, strongly set with differential privacy when record linkage is present, and Xuan Bi on a fully decentralised federated learning with local exchanges of global gradients that limit privacy leaks. (With a mention of gossip learning I hadn’t seen previously!)

I enjoyed even more the session due to Naisyin Wang giving a discussion on the three talks, in  closes the particular because I had not seen her in years, if a few times since she was my teaching assistant in Cornell in a Bayesian decision theory class I was building on the spot (with Linda Zhao as a student!). Which nicely closes the loop given the topic of my talk, of which she was quite supportive! While calling for debiasing post-processing and wondering about a two-dimensional decision theoretic perspective rather than a unidimensional one under hard privacy constraints.

On the way out, after parting from the few friends remaining in the convention centre, I spotted the nearby (steel) bridge being raised, although I could not see the boat responsible for it. I had been unaware of this possibility while running over and under it, as well as swimming thrice under it. And we left Portland in the early afternoon, heading for Seattle and a celebration of Adrian Raftery’s career.

JSM 2024, Portland, Day 3

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on August 9, 2024 by xi'an

Bayesian contributed session as the first round of the third day (with a choice of five parallel sessions featuring Bayesian topics!!, actually easier to pick than among the following eight parallel sessions of the 10:30 schedule!!!), with a talk by Tahir Ekin on adversarial outlier detection that could connect with our Oceaner(c) privacy concerns. Then one involving spike & slab (a theme to figure prominently in this special day!!) in mixed response models by Sameer Deshpande, seeking a (unBayesian!) MAP for a latent variable model by Monte Carlo EM. Followed by a talk by Yunyi Shen on completely random measures for estimating the (distribution of the) number of species in heterogeneous populations. Next, Valentin Zulj on (frequentist rather than) Bayesian stacking, on estimating optimal weights for model averaging (which should be posterior probabilities in a pure Bayesian mindframe), including a score function that could lead to generalised Bayesian inference on said weights. Finishing with a talk by Chaegeun Song on correcting Bayesian credible sets towards (frequentist, again!!!) exact coverage for classification (which reminded me of my very first paper with George on correcting frequentist confidence for Binomial observations). With which I could not really engage as seeking a specific coverage level did not seem relevant, imho, but I appreciated the wheel plot representation.My second morn session was about modern (what else?!) sampling algorithms, although I spent the first dozen minutes wondering whether or not I had entered the wrong room. Until Tianhao Wang focussed on Thompson sampling for bandits. It did prove far enough from my interest for my (sleep deprived) attention to drift too quickly. Only the talk by Yuchen Wu on a spike & slab (as suits the day!) challenge captured enough this wandering attention. Crossing further into my realm of primary topics by considering a target distribution that is a product of distributions. But I did not get from her presentation how a product measure decomposition was inducing higher efficiency (and did not find answers within the arXived preprint). Unless it exploited specific features of the target, like conditional independence between the components. The last talk was by Brice Huang on sampling low temperature Gibbs measures using stochastic localisation.

After coming upon a row of food trucks across the conference centre and being unfairly attracted by an Ethiopian injera picture into a terrible wrap, I returned for the Skeptical about AI session, just a few minutes late, only to find accessing the session was impossible! Quite sad to miss the presentations and the arguments (even though I had heard a previous talk by Genevera Allen when visiting Rutgers two years ago). As a second best, I then joined the recent (of course!) Advances in Bayesian Computation (aka ABC?!) session with a medley of topics, including a data subset versus data sketching model reduction by Sudipto Saha. Which could have consequences on our privacy strategies. And marginal evidence estimation for the Bayesian Lasso by Christopher Hans while avoiding data completion. And another latent variable model with a sequential variational Bayes approach by Bao Anh Vu, using at one point Cappé et al. (2005) EM-based approximation to the log likelihood gradient. Finishing by a back-to-the-future talk by Luke Duttweiler on MCMC convergence diagnostics. Comparing several chains via proximity maps that themselves require some preliminary knowledge about the MCMC kernel. (Nice title though, “the traceplot thickens”!)The crux of the day was however the 2024 COPSS Award ceremony with several friends featuring among the recipients, Danielle Durante for the Emerging Leaders Award, Regina Liu for the Elizabeth L. Scott Award and Veronika Rockova for the Presidents’ Award. Congrats!!!



Bayesian Nonparametrics Networking Workshop 2023

Posted in Statistics, Travel, University life with tags , , , , , , , , on December 3, 2023 by xi'an

Natural statistical science [#2]

Posted in Statistics with tags , , , , , , , , , , , , , , , on November 23, 2023 by xi'an

A rare occurrence of a Bayesian statistics paper in Nature with this “State estimation of a physical system with unknown governing equations” by Course and Nair. A variational Bayes modelling of a state system observed with noise, but without a physical model on the state (SDE) evolution itself. Which means a prior is set on a non-parametric or neural representation of the drift and a linear approximation is used for the variational approximation, leading to a Gaussian process as the approximate distribution. While this applies to highly complex models, like orbiting black holes, it is somewhat a surprise to meet this application of variational inference in a prestigious general science journal like Nature. (The picture above was taken on the train from Marseille at the end of the Bayes Fall school.)

“The approach is based on a technique called Bayesian inference, which is used widely, but which can be computationally challenging for complex systems.” B. Keith

Familial inference

Posted in Statistics, University life with tags , , , , , , , , , on October 3, 2023 by xi'an

An ISBA-BNP webinar on Wednesday, 4 October, at 17:00 UTC by my friend Steve McEachern:

Familial inference: Tests for hypotheses on a family of centers

Many scientific disciplines face a replicability crisis. While these crises have many drivers, we focus on one. Statistical hypotheses are translations of scientific hypotheses into statements about one or more distributions. The most basic tests focus on the centers of the distributions. Such tests implicitly assume a specific center, e.g., the mean or the median. Yet, scientific hypotheses do not always specify a particular center. This ambiguity leaves a gap between scientific theory and statistical practice that can lead to rejection of a true null. The gap is compounded when we consider deficiencies in the formal statistical model. Rather than testing a single center, we propose testing a family of plausible centers, such as those induced by the Huber loss function (the Huber family). Each center in the family generates a point null hypothesis and the resulting family of hypotheses constitutes a familial null hypothesis. A Bayesian nonparametric procedure is devised to test the familial null. Implementation for the Huber family is facilitated by a novel pathwise optimization routine. Along the way, we visit the question of what it means to be the center of a distribution. The favorable properties of the new test are demonstrated theoretically and in case studies.
This is joint work with Ryan Thompson (University of New South Wales), Catherine Forbes (Monash University), and Mario Peruggia (The Ohio State University).