Our Warwick PhD student Shreya Sinha-Roy—who is now looking for a postdoctoral position next semester!—, along with Sherman Khoo, Ritabrata Dutta and myself, has now completed a paper on shrinkage priors for implicit generative models. That is, models based on deep neural networks and hence associated with intractable likelihoods. The work centres on developing and assessing an efficient training mechanism for these models, leveraging on tools from Bayesian model averaging using shrinkage (yay!) priors inspired from Lasso (rather than from my PhD years!) and generalized Bayes. In this large p (dimension of parameters) and small n (sample size of data) scenario, those sparsity inducing priors have been successfully used for linear regression when p is much larger than n, but have not been applied to implicit generative models due to the intractability of the likelihood function of the parameters of the model given observed data. Adapting a scoring rule posterior based on a strictly proper scoring rule as in generalized Bayes, we propose a block SGMCMC within Gibbs sampling mechanism to handle high dimensional parameter space for learning a sparse Bayesian model averaged neural implicit generative model in a sample efficient way. We illustrate excellent performance of our proposed method for p (much larger than n) linear regressions and three applications of neural generative models in tasks relevant to weather forecasting to reinforcement learning
Archive for shrinkage
back to shrinkage!
Posted in Books, Statistics, University life with tags Bayesian lasso, generalised Bayesian inference, generative model, Gibbs sampling, intractable likelihood, MathPhDInFrance, Rouen, shrinkage, Université de Rouen, University of Warwick, Warwickshire on February 12, 2026 by xi'ana (sunny, crisp) day at ICSDS 2025
Posted in pictures, Running, Statistics, Travel, University life with tags 5thSYSORM, AI, Andalucía, Bayesian nonparametrics, Bill Strawderman, conferences, CRiSM, Croatia, deviance information criterion, DIC, ellipsoid, generalised Bayesian inference, Guadalquivir, ICSDS 2025, ICSDS 2026, IMS, martingale posterior, minimaxity, prediction, Réal Fabrica deTabacos de Sevilla, Sevilla, shrinkage, shrinkage estimation, Spain, Split, transformer, Universidad de Sevilla, University of Warwick, urn on December 19, 2025 by xi'an
While my first day at ICSDS 2025 was somewhat hectic, having realised late the night before that I was giving a talk!—I had forgotten I had submitted a title at registration time and never received any communication from the organisers, including (or excluding) a request for an abstract. I thus hastily updated my November talk in Sevilla for my December talk in Sevilla! but paid less attention than needed to the sessions I attended—, Wednesday was more peaceful—esp. after a 16K run along the Guadalquivir—and I engaged into two great Bayesian learning sessions, one that seemed designed for me!, involving my (40y long friend) Ed George on his latest result on proper prior minimaxity and shrinkage, with our late friend Bill Strawderman as a co-author since they worked on the problem prior to Bill’s demise, Charles Margossian on variational inference preserving some symmetries in the target and hence keeping the same statistics, with elliptically symmetric families, and Fletcher Christensen on DIC for some mixed models, with references to our “DIC’s eights” paper (but still picking one version of DIC in the end!)
The second session was on prediction learning!—with me as the chair, as I realized one minute before! AI !—with (my friend) Veronika Rockova using AI predictions as a prior predictive and connecting them with Bayesian nonparametrics, Kenyon Ng (who visited me last Spring) on a similar approach using pretrained transformers like TabPFN and martingale posterior inference, Lorenzo Cappello in a generalisation of martingale prediction and Andrea Ghiglietti on the mathematics of an involved urn system.

The afternoon session was a plenary talk by Daniela Witten in the magnificent building of the Real Fabrica de Tabacos, but the room was unfortunately too small for the audience and I could not enter. Hopefully her talk will have a significant intersection with the CRiSM colloquium she delivers in Warwick late January. I thus walked around the old town till the following poster session, held in the Real Fabrica courtyard, under the sun. As I got involved into a deep discussion of the relevance of mirror meetings (which I defend!) versus the dangers on principal (parent) conferences (which can be mitigated by the mirror conference participants registering, to some extent, for the principle one)—more to come on the ‘Og and in the ISBA Bulletin!—, I did not peruse the available posters, sorry…

And, by the way, the conference organisers also revealed the location of ICSDS 2026 which is Croatia, my first bet! In the city of Split we visited in 2023.
BayesComp 2025.4
Posted in pictures, Running, Statistics, Travel, University life with tags ABC, ABC model selection, Adam, approximate Bayesian inference, BayesComp 2025, Bayesian GANs, Bayesian lasso, Bayesian neural networks, Bayesian optimisation, Bayesian paradigm, Bayesian predictive, Bayesian robustness, Bayesian semi-parametrics, Baysian learning, BIC, chili crab, differential privacy, harmonic mean estimator, homomorphic encryption, hot pot, Laplace approximation, Les Houches, maximum mean discrepancy, mee siam, mixture estimation, National University Singapore, NUS, Peranakan cuisine, plenary speaker, power posterior, privacy laws, random kernel MCMC, RATP, RER, Roberta, safe Bayes, sequential importance sampling, shrinkage, shrinkage estimation, simulation-based inference, Singapore, SNCF, splines, Stein divergence, stochastic gradient MCMC, stochastic optimisation, summary statistics, Swendsen-Wang algorithm, Sylvia Frühwirth-Schnatter, Szechuan cuisine, treadmill, University of Warwick, unknown number of components, variational Bayes methods, Wasserstein distance, William Strawderman, WU Wirtschaftsuniversität Wien, zigzag algorithm on June 21, 2025 by xi'an
The third and final day of the (main) conference started tih Emtiyaz Khan’s plenary talk on adaptive Bayesian intelligence. Or, imho, [adaptive [Bayesian]] intelligence, with the brackets indicating redundancy since intelligence need include adaptivity and [intelligent] adaptivity need proceed in a Bayesian way! Focussing first on the Bayesian learning rule via variational Bayes (with a stress on Kingma’s 1994 Adam optimisation algorithm, the “most cited paper” [in machine learning]) where learning boils down to gradient steps (due to the exponential family structure), themselves versions of Taylor (or Laplace) approximations). With an interesting vision of Bayesian updating as accounting for prediction mismatch. (I missed the connection Roberta in IMDb appearing in one slide!)
The following session offered no dilemma [sorry, Alex, Axel, Chris, Robert, Sumeet, Victor!] since it included the federated learning session I organised, with Louis Asslet, Conor Hassan, and Jean-Michel Marin as speakers. Louis’ talk was on confidential [homomorphic] accept-reject algorithms to learn from other sources, while preserving (differential?) privacy, part of which came during Les Houches workshops I organised this Spring and the one before. Exploiting the additive features of log-likelihoods and exponential variates and adopting a testing perspective on privacy. Conor motivated his model with the Australian cancer atlas project Kerrie Mengersen and others have been developing over the years. The federated approach relies on variational approximations that return the same answer as an exact resolution, but more efficiently. (From a privacy perspective, I wonder at the impact of variational approximations on protecting the data, which boils down to a choice of (sufficient) statistics for the exponential families behind those approximations.) For more complicated models incorporating spatial dependence prohibits full Bayesian inference, unfortunately. Jean-Michel commented on the richness of methods for simulation-based inference, incl. model choice. His focus was on using sequential neural likelihood estimation and sequential importance sampling to approximate evidence. As in the Read Paper of Del Moral et al. (2006). Mentioning a neural version of the harmonic mean estimator by Spurio Mancini et al. (2023)! I wondered at the degree of (Rao-Blackwell) recycling involved in the computation, Jean-Michel’s answer being that AMIS is soon coming [in a theatre near you!].

The afternoon sessions did offer any reprieve in the choice of topic! I first went to Approximate Methods for Accelerated Sampling, with Rong Tang evaluating the informativeness of summary statistics through a divergence evaluation. Using autoencoders to replace the intractable posterior, with sliced minimal model discrepancy (MMD) and (pseudo?) score matching loss for divergences (reminding me of indirect inference and synthetic likelihood). Yun Yang discussed a variational proposal to estimate the number of components in a mixture model. Surprising given the multimodal structure of mixture posteriors. And the overall irregularity of (evil!) mixture models. But I could not figure out from the talk the form of the approximation.

On the food scene, tasted a nice and spicy Peranakan rice vermicelli dish called Mee Siam yesterday in a campus restaurant, which sustained me fore the rest of the day, including the ABC s/webinar. And another spicy hot pot today at NUS, to catch up on veggies, while missing the chili crab local specialty on that trip.
BayesComp 2025.3
Posted in pictures, Running, Statistics, Travel, University life with tags abalone, ABC, ABC model selection, BayesComp 2025, Bayesian GANs, Bayesian lasso, Bayesian neural networks, Bayesian predictive, Bayesian robustness, Bayesian semi-parametrics, changepoint detection, equator, Gibbs posterior, Henri Poincaré, horseshoe prior, humidity, Kingman's coalescent, Langevin MCMC algorithm, local regression, National University Singapore, NUS, Ocean, OWABI, particle filters, PDMP, plenary speaker, power posterior, random kernel MCMC, RATP, RER, safe Bayes, shrinkage, shrinkage estimation, Singapore, SNCF, splines, Stein divergence, stochastic gradient MCMC, Swendsen-Wang algorithm, Sylvia Frühwirth-Schnatter, Szechuan cuisine, treadmill, University of Warwick, Wasserstein distance, William Strawderman, WU Wirtschaftsuniversität Wien, zigzag algorithm on June 20, 2025 by xi'an
The second day of the conference started with a cooler and less humid weather (although this did not last!), although my brain felt a wee bit foggy from a lack of sleep (and I almost crashed while running on the hotel treadmill, at 14.5km/h!), and the plenary talk of my friend of many years Sylvia Früwirth-Schnatter on horseshoe priors and time-varying time series (à la West). With a nice closed-form representation involving hypergeometric functions of the second kind (my favourite!), with the addition of a triple-Gamma prior. Sylvia stressed on the enormous impact of the prior choice on change-point detection, which was already the point in the original horseshoe paper (as opposed to George’s Lasso prior). Without incorporating any specific modelling on potential change-point, fair enough given that the parameter is moving with time, unhindered. Her MCMC choices involved discrete parameters with Negative Binomial and Poisson parameters, allowing for partially integrated or collapsed solutions. Possibly further improved by Swendsen-Wang steps.

I then attended the (advanced) Langevin session after agonising upon my choice for a wealth of options! Sam Power presented a talk linking simulation with optimisation targets, over measure spaces. With Wasserstein gradient flow algorithms that resemble Langevin algorithms once discretised by a particle system. (A natural resolution producing a somewhat unnatural form of measure estimator since made of Dirac masses, from which very little can be learned.) Then [my Warwick colleague & coauthor] Any Wang on underdamped Langevin diffusions. when Poincaré‘s inequality fails, but convergence (in total variation) still occurs. Followed by Peter Whalley on splitting methods (where random hypergeometric subsampling dominates Robbins-Monro) and stochastic gradient algorithms, in a connected (to the previous talks) way since involving underdamped aspects. (With a personal discovery of Polyak’s heavy ball method.)

The afternoon session saw me facing a terrible dilemma with three close friends talking at the same time! Eventually opting for PDMPs, over simulation-based inference and recalibration for approximate Bayesian methods. Kengo Kamatani gave a general introduction to PDMPs, before explaining the automated implementation he considered with Charly Andral (during Charly’s visit to ISM, Tokyo, two summers ago). Towards accelerating the generation of the jump time. Then Luke Hardcastle applied PDMPs for survival prediction, using spike & slab priors and sticky PDMPs. And Jere Koskela (formerly Warwick) extended zig-zag sampling to discrete settings (incl. Kingman’s coalescent.)
The (rather long) day was not over yet since we had planned an extra on-site OWABI seminar & webinar with two participants in the conference, Filippo Pagani (Warwick and OCEAN postdoc) using fusion for federated learning, with a trapezoidal approximation, and Maurizio Filippone on GANs as hidden perfect ABC model selection, a GAN providing an automatic density estimator… With astounding Gemini-generated cartoons! Videos are soon to be available. A big congrats to the speakers who managed to convey their ideas and results despite the late hour! (On the extra-academic side, I was invited last night to a genuine Szechuan dinner in Chinatown, with a large array of spicy dishes if not that spicy!, and a rare opportunity to taste abalone. And bullfrogs. Quite a treat! And a good reason to skip dinner altogether!)

thou shalt not slice thine spaghetti
Posted in Statistics with tags geodesics, manifold, pasta machine, reversibility, Riemann manifold, Riemannian measure, seminar, shrinkage, slice sampler, spaghetti, University of Warwick on May 1, 2024 by xi'an
This 2023 work on Slice sampler on manifolds, as presented in the algorithms seminar in Warwick during a recent visit of by Mareike Hasenpflug, consists in designing and validating slice samplers for distributions on manifolds. It is mildly connected to some current work on MCMC algorithms on manifolds through coupling techniques by [my friends & coauthors] Elena Bortolado, Pierre Jacob, and Robin Ryder (who escape temporarily the manifold at each step). As in Neal (2003), uniform draws from the (super)level sets are replaced there with one-step Markov moves within the level set, that is, slice sampler moves. The slice sampler actually generalises Neal’s (2003) stepping-out and shrinkage steps rather closely. Based on the standard notion of the Riemannian measure induced by the very structure of the manifold, the model therein assumes that the simulation target is available as a closed-form if unormalised density p(x) against that measure, meaning that problems where the distribution is a push-forward one induced by a mapping onto the manifold are not necessary manageable. The slide sampler is decomposed into choosing (1-dimensional) geodesics defined by the manifold (and generalising great circles), uniformly, and then sampling by this one-dimensional slice sampling over the geodesic, under the level set constraint. Meaning that those geodesics must be manageable enough. (Note that the concept of stepping-out does not mean that the chain ever escapes from the manifold.) Demonstrating the validity and reversibility proves a challenging task.