Our Warwick PhD student Shreya Sinha-Roy—who is now looking for a postdoctoral position next semester!—, along with Sherman Khoo, Ritabrata Dutta and myself, has now completed a paper on shrinkage priors for implicit generative models. That is, models based on deep neural networks and hence associated with intractable likelihoods. The work centres on developing and assessing an efficient training mechanism for these models, leveraging on tools from Bayesian model averaging using shrinkage (yay!) priors inspired from Lasso (rather than from my PhD years!) and generalized Bayes. In this large p (dimension of parameters) and small n (sample size of data) scenario, those sparsity inducing priors have been successfully used for linear regression when p is much larger than n, but have not been applied to implicit generative models due to the intractability of the likelihood function of the parameters of the model given observed data. Adapting a scoring rule posterior based on a strictly proper scoring rule as in generalized Bayes, we propose a block SGMCMC within Gibbs sampling mechanism to handle high dimensional parameter space for learning a sparse Bayesian model averaged neural implicit generative model in a sample efficient way. We illustrate excellent performance of our proposed method for p (much larger than n) linear regressions and three applications of neural generative models in tasks relevant to weather forecasting to reinforcement learning
Archive for Bayesian lasso
back to shrinkage!
Posted in Books, Statistics, University life with tags Bayesian lasso, generalised Bayesian inference, generative model, Gibbs sampling, intractable likelihood, MathPhDInFrance, Rouen, shrinkage, Université de Rouen, University of Warwick, Warwickshire on February 12, 2026 by xi'anBayesComp 2025.4
Posted in pictures, Running, Statistics, Travel, University life with tags ABC, ABC model selection, Adam, approximate Bayesian inference, BayesComp 2025, Bayesian GANs, Bayesian lasso, Bayesian neural networks, Bayesian optimisation, Bayesian paradigm, Bayesian predictive, Bayesian robustness, Bayesian semi-parametrics, Baysian learning, BIC, chili crab, differential privacy, harmonic mean estimator, homomorphic encryption, hot pot, Laplace approximation, Les Houches, maximum mean discrepancy, mee siam, mixture estimation, National University Singapore, NUS, Peranakan cuisine, plenary speaker, power posterior, privacy laws, random kernel MCMC, RATP, RER, Roberta, safe Bayes, sequential importance sampling, shrinkage, shrinkage estimation, simulation-based inference, Singapore, SNCF, splines, Stein divergence, stochastic gradient MCMC, stochastic optimisation, summary statistics, Swendsen-Wang algorithm, Sylvia Frühwirth-Schnatter, Szechuan cuisine, treadmill, University of Warwick, unknown number of components, variational Bayes methods, Wasserstein distance, William Strawderman, WU Wirtschaftsuniversität Wien, zigzag algorithm on June 21, 2025 by xi'an
The third and final day of the (main) conference started tih Emtiyaz Khan’s plenary talk on adaptive Bayesian intelligence. Or, imho, [adaptive [Bayesian]] intelligence, with the brackets indicating redundancy since intelligence need include adaptivity and [intelligent] adaptivity need proceed in a Bayesian way! Focussing first on the Bayesian learning rule via variational Bayes (with a stress on Kingma’s 1994 Adam optimisation algorithm, the “most cited paper” [in machine learning]) where learning boils down to gradient steps (due to the exponential family structure), themselves versions of Taylor (or Laplace) approximations). With an interesting vision of Bayesian updating as accounting for prediction mismatch. (I missed the connection Roberta in IMDb appearing in one slide!)
The following session offered no dilemma [sorry, Alex, Axel, Chris, Robert, Sumeet, Victor!] since it included the federated learning session I organised, with Louis Asslet, Conor Hassan, and Jean-Michel Marin as speakers. Louis’ talk was on confidential [homomorphic] accept-reject algorithms to learn from other sources, while preserving (differential?) privacy, part of which came during Les Houches workshops I organised this Spring and the one before. Exploiting the additive features of log-likelihoods and exponential variates and adopting a testing perspective on privacy. Conor motivated his model with the Australian cancer atlas project Kerrie Mengersen and others have been developing over the years. The federated approach relies on variational approximations that return the same answer as an exact resolution, but more efficiently. (From a privacy perspective, I wonder at the impact of variational approximations on protecting the data, which boils down to a choice of (sufficient) statistics for the exponential families behind those approximations.) For more complicated models incorporating spatial dependence prohibits full Bayesian inference, unfortunately. Jean-Michel commented on the richness of methods for simulation-based inference, incl. model choice. His focus was on using sequential neural likelihood estimation and sequential importance sampling to approximate evidence. As in the Read Paper of Del Moral et al. (2006). Mentioning a neural version of the harmonic mean estimator by Spurio Mancini et al. (2023)! I wondered at the degree of (Rao-Blackwell) recycling involved in the computation, Jean-Michel’s answer being that AMIS is soon coming [in a theatre near you!].

The afternoon sessions did offer any reprieve in the choice of topic! I first went to Approximate Methods for Accelerated Sampling, with Rong Tang evaluating the informativeness of summary statistics through a divergence evaluation. Using autoencoders to replace the intractable posterior, with sliced minimal model discrepancy (MMD) and (pseudo?) score matching loss for divergences (reminding me of indirect inference and synthetic likelihood). Yun Yang discussed a variational proposal to estimate the number of components in a mixture model. Surprising given the multimodal structure of mixture posteriors. And the overall irregularity of (evil!) mixture models. But I could not figure out from the talk the form of the approximation.

On the food scene, tasted a nice and spicy Peranakan rice vermicelli dish called Mee Siam yesterday in a campus restaurant, which sustained me fore the rest of the day, including the ABC s/webinar. And another spicy hot pot today at NUS, to catch up on veggies, while missing the chili crab local specialty on that trip.
BayesComp 2025.3
Posted in pictures, Running, Statistics, Travel, University life with tags abalone, ABC, ABC model selection, BayesComp 2025, Bayesian GANs, Bayesian lasso, Bayesian neural networks, Bayesian predictive, Bayesian robustness, Bayesian semi-parametrics, changepoint detection, equator, Gibbs posterior, Henri Poincaré, horseshoe prior, humidity, Kingman's coalescent, Langevin MCMC algorithm, local regression, National University Singapore, NUS, Ocean, OWABI, particle filters, PDMP, plenary speaker, power posterior, random kernel MCMC, RATP, RER, safe Bayes, shrinkage, shrinkage estimation, Singapore, SNCF, splines, Stein divergence, stochastic gradient MCMC, Swendsen-Wang algorithm, Sylvia Frühwirth-Schnatter, Szechuan cuisine, treadmill, University of Warwick, Wasserstein distance, William Strawderman, WU Wirtschaftsuniversität Wien, zigzag algorithm on June 20, 2025 by xi'an
The second day of the conference started with a cooler and less humid weather (although this did not last!), although my brain felt a wee bit foggy from a lack of sleep (and I almost crashed while running on the hotel treadmill, at 14.5km/h!), and the plenary talk of my friend of many years Sylvia Früwirth-Schnatter on horseshoe priors and time-varying time series (à la West). With a nice closed-form representation involving hypergeometric functions of the second kind (my favourite!), with the addition of a triple-Gamma prior. Sylvia stressed on the enormous impact of the prior choice on change-point detection, which was already the point in the original horseshoe paper (as opposed to George’s Lasso prior). Without incorporating any specific modelling on potential change-point, fair enough given that the parameter is moving with time, unhindered. Her MCMC choices involved discrete parameters with Negative Binomial and Poisson parameters, allowing for partially integrated or collapsed solutions. Possibly further improved by Swendsen-Wang steps.

I then attended the (advanced) Langevin session after agonising upon my choice for a wealth of options! Sam Power presented a talk linking simulation with optimisation targets, over measure spaces. With Wasserstein gradient flow algorithms that resemble Langevin algorithms once discretised by a particle system. (A natural resolution producing a somewhat unnatural form of measure estimator since made of Dirac masses, from which very little can be learned.) Then [my Warwick colleague & coauthor] Any Wang on underdamped Langevin diffusions. when Poincaré‘s inequality fails, but convergence (in total variation) still occurs. Followed by Peter Whalley on splitting methods (where random hypergeometric subsampling dominates Robbins-Monro) and stochastic gradient algorithms, in a connected (to the previous talks) way since involving underdamped aspects. (With a personal discovery of Polyak’s heavy ball method.)

The afternoon session saw me facing a terrible dilemma with three close friends talking at the same time! Eventually opting for PDMPs, over simulation-based inference and recalibration for approximate Bayesian methods. Kengo Kamatani gave a general introduction to PDMPs, before explaining the automated implementation he considered with Charly Andral (during Charly’s visit to ISM, Tokyo, two summers ago). Towards accelerating the generation of the jump time. Then Luke Hardcastle applied PDMPs for survival prediction, using spike & slab priors and sticky PDMPs. And Jere Koskela (formerly Warwick) extended zig-zag sampling to discrete settings (incl. Kingman’s coalescent.)
The (rather long) day was not over yet since we had planned an extra on-site OWABI seminar & webinar with two participants in the conference, Filippo Pagani (Warwick and OCEAN postdoc) using fusion for federated learning, with a trapezoidal approximation, and Maurizio Filippone on GANs as hidden perfect ABC model selection, a GAN providing an automatic density estimator… With astounding Gemini-generated cartoons! Videos are soon to be available. A big congrats to the speakers who managed to convey their ideas and results despite the late hour! (On the extra-academic side, I was invited last night to a genuine Szechuan dinner in Chinatown, with a large array of spicy dishes if not that spicy!, and a rare opportunity to taste abalone. And bullfrogs. Quite a treat! And a good reason to skip dinner altogether!)

Seminal ideas and controversies in Statistics [book review]
Posted in Books, Mountains, pictures, Statistics, Travel, University life with tags ABC model choice, ACEMS, Adelaide, anti-vaccine, ASA, ASC 2012, Australia, autism, Bayesian bootstrap, Bayesian lasso, Biometrika, blonde, book cover, bootstrap, Bruce Lindsay, Calyampudi Radhakrishna Rao, Census, CHANCE, classics, controversies, CRC Press, Cuillin ridge, David Blackwell, econometrics, EDA, EM algorithm, estimating equations, FDR, finite mixtures, Gibbs sampling, Glasgow, Grace Wabha, Heinz, history of statistics, hypothesis testing, ICMS, International Prize in Statistics, James-Stein estimator, JASA, Jerzy Neyman, John Tukey, JRSS, machine learning, mathematical statistics, MCMC, Neyman-Pearson tests, pickles, posterior predictive check, Rashomon, replicability, ridge regression, Roderick Little, Ronald Fisher, Ryûnosuke Akutagawa, sabbatical, Scotland, shrinkage estimation, Skye, Stein's paradox, Steve Fienberg, the inaccessible pinacle, The Likelihood Principle, toothbrush moustache, Université Paris Dauphine, World Fertility Survey on May 24, 2025 by xi'an
CRC Press sent CHANCE this book for review. Since the topic was of clear interest to me, with an author who significantly contributed to the field—my only recollection meeting Roderick Little was during the Australian Statistical Conference in Adelaïde, in 2012, at the start of my Oz 2012 Tour!—, I took the opportunity of the nearest weekend to browse through Seminal ideas and controversies in Statistics. I like very much the idea of selecting a dozen key papers in the history of Statistics and of discussing why. In fact, this reminded me of my classics seminar, which lasted the few years I was 100% in charge of the Master program in Dauphine (and which I hope I could restart!). Checking the list of the papers I then suggested my students, I see some overlap with 9 papers out of the 15 groups. (I also remember Steve Fienberg making suggestions for that list, while he was spending a sabbatical in Paris at CREST.) Given that community of focus and purpose, and contrary to my wont, I have really very little of substance to criticize or wish about the book. The less when reading the following
“On a personal note, I met Yates [author of a 1984 paper on tests for 2×2 contingency tables discussing the relevance of conditioning on one or both margins], a charming man, when I was a young graduate student who knew next to nothing about statistics; we discussed the joys of traversing the Cuillin Ridge in Skye.”
since completing that ridge remains high in my mountain-climbing bucket-list! (Possibly next year, since we are running an ICMS workshop on the Island.)
The first paper in the series is more than a foundational paper since (The) Fisher’s 1922 paper is about creating (almost) ex nihilo the field of (modern) mathematical statistics. I don’t know if there is any equivalence in other scientific disciplines of such an impact (and of such a man)… Roderick Little manages to convincingly engage with Fisher’s dismissive views on (not yet called) Bayesian analysis, although, to the latter’s defence, the formalisation of Bayesian inference at that time had not yet emerged. The second chapter is discussing Yates’ 1984 paper on tests for 2×2 contingency tables that he wrote 50 years after writing the original one in the first volume of JRSS. Roderick Little adds a detailed Bayesian analysis with the three standard reference priors, Jeffreys’ version proving quite close to Fisher’s exact test (conditional on both margins). The third chapter is aiming at the generic challenge of hypothesis testing, from the well-known opposition between Fisher and Neyman (both on the cover), to questioning the sanity of hard-set thresholds (with a mention of our American Statistician call to abandon (shi)p!). The later (thus) refers to the recent literature on the replicability crisis and the now famous ASA statement on p-values by Ron Wasserstein and Nicole Lazar, analysed in the chapter. But I would have like to read another full section on alternatives to hypothesis testing. While now a niche interest (imho), Fisher’s attempt at creating a posterior distribution without a prior, aka fiducial inference, is discussed in Chapter 4 with the Behrens-Fisher problem as the illustrating example. The chapter feels rather anticlimactic, with the comparison relying on the (Malay) Ghosh and Kim (2001) simulation results.
Birnbaum’s (1962) likelihood principle is the topic of Chapter 5 (and I cannot remember any of my students choosing this paper over the years, although there was at least one). Roderick Little recalls some sentences from the JASA discussion as an appetiser, a reminder of the time when these discussions could turn in scathing attacks. The chapter contains excerpts from Berger and Wolpert (1988)—which they were writing while I was spending a year at Purdue and which I have always recommended to my PhD students, albeit not for the classic seminar. It then moves to the controversies that surround this principle since its inception, in particular those accumulated by Deborah Mayo (also on the cover) as reported on the ‘Og. In the recent years, I have become less excited about the LP, in part due to the imprecision in its statement, which opens the door to conflicting interpretations. And in part due to the scarcity of models with non-trivial sufficient statistics. (I am also wondering if the sufficiency issue we highlighted in our ABC model choice criticism does relate to the mixture example at the end of the chapter.)
The next chapter is one all for compromise, through the calibrated Bayes perspective that credible statements should be close to confidence statements in the long run. Which I remember him presenting at ASC 2012. The concept is found in the very 1984 paper by Don Rubin (also on the cover) that contains the concept behind Approximate Bayesian Computation (ABC). And the chapter proceeds by listing strengths and weaknesses of frequentist and Bayesian perspectives, towards a fusion of both., e.g. though posterior predictive checks.
While the choice of a (general public) paper from Scientific American may sound surprising in Chapter 7, with Efron’s (on the cover) and Morris’ 1977 Stein’s paradox, I cannot but applaud, the more because this was the first paper I read when starting my PhD on the James-Stein estimators. Although this may sound like happening eons ago, the James and Stein (1961) paper—which is my age!—”created a considerable backlash” by toppling unbiasedness from its pedestal and exhibiting a paradox that 1+1+1≠3… Which Little reinterprets via a random effect (or Bayesian hierarchical) model. (And a chapter where I learned that Little’s father was a journalist, a characteristic he shared with Bruce Lindsay, as I found at Blonde, Glasgow, during an ICMS workshop). Relatedly, the next chapter is about the “57 varieties [of regression] paper” by Demptster, Schatzoff and Wermuth (1977). Apparently connected with Heinz 57 varieties of pickles. The paper considers Stein and ridge and variable selections versions for variable selection. The chapter also covers (Bayesian) Lasso and BART, as well as a brief all too brief mention of Spike & Slab priors—with my friend Veronika Ročková missing from the authors’ index!—, but I was expecting from the title other, robust, forms of regression like L¹ regression and econometrics digressions. Chapter 10 can however been seen as a proxy since covering generalized estimating equations from a 1986 Biometrika paper of Liang and Zeger, with no Bayesian aspect (and an expected appearance of Communications in Statistics B).
Chapter 9 covers the almost immediately classic 1995 paper of Benjamini and Hochbeg on multiple regressions (that Series B turned into a discussion paper ten years later!). Although it spends more time on Berry’s (2012) recommendations than on FDR. The computational Chapter 11 brings together Efron’s (1979) bootstrap [with his picture on the cover] and MCMC, represented by the founding paper of Gelfand and Smith (1990, if mistakenly set in 1988 on p140). A bit of a strange mix imho as the former is more inferential than computational. And not giving the EM algorithm that much space. And not questioning MCMC methods as a good proxy to posterior distributions. Tukey’s Future of Data Analysis (as founding exploratory data analysis) and Breiman’s Two cultures (as launching statistical machine learning) meet in Chapter 12. (With a reminder that the latter invokes Occam’s razor—which may not be that appropriate for hugely overparameterised machine learning black boxes—and…the Rashomon principle! Meaning that distinct models may all fit the same data. Let me nitpickingly add the reference to Ryûnosuke Akutagawa as the author of Rashômon and other stories that Kurosawa adapted in his splendid movie). The chapter contains critical remarks from David Cox, Brad Efron, David Bickel, and Andrew Gelman, with a further section on Little’s view on modelling.
The last three chapters are on design and sampling, in connection with Little’s (and Rubin’s) works in the area. With a 1934 paper of Neyman (whose picture on the cover could have been chosen differently, albeit no fault of Neyman [or of Little!] that his toothbrush style of moustache dramatically got out of fashion!). With a return to calibrated Bayes and a reminiscence of Little’s time at the World Fertility Survey but (apparently) no mention of the probabilistic aspects of modern censuses (that saw my friends Steve Fienberg on the one side and Larry Brown and Marty Wells on the other side argue for and against it!), again relating to the reliance on statistical models. Chapter 14 relates randomized clinical trials to causality, which makes a (worthy) appearance there. Roderick Little also makes a clear case there against the retracted study linking vaccines and autism, a call that will unlikely not reach the current Trump administration and its Secretary of Health.
The book concludes with a list of twenty style and grammar suggestions for improved writing.
As should be crystal-clear from the above, I quite enjoyed the book and would definitely use its reading list in a graduate course whenever the opportunity arises. Once again, some choices are more personal to the author than others, and I would have place more emphasis on the fantastic Dawid, Stone and Zidek (1973)—with Jim Zidek also missing from the author index—, but all make sense in a walk through statistical classics. Let me however regret the absence therein of major actors like, e.g., D. Blackwell, C.R. Rao, or G. Wahba (except in a stylistic example p199), two of whom were awarded the International Prize in Statistics.
[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE.]
IISA 2024
Posted in Statistics, Travel, University life with tags Bayes factor, Bayesian bootstrap, Bayesian computational methods, Bayesian GANs, Bayesian lasso, Bayesian nonparametrics, betel nut, BNP, Cochin University of Science and Technology, difference in differences model, EM algorithm, IISA 2024, India, insufficient statistic, ISI, Kerala, Kochi, privacy, random walk, Wasserstein distance on January 12, 2025 by xi'an
The IISA 2024 conference was held at the Cochin University of Science and Technology (CUSAT), where talks and posters took place. Among the sessions I attended, I attended a talk on colliding random walks that made three or more impossible in the limit, missed the only privacy talk that I could have attended by Vinayak Rao for indulging into a dawnish swimming session in the very decent hotel pool. In the first session I organised, loosely connected with BNP, Antonietta Mira presented a recent work on predicting EU carbon compensation rates, using a information imbalance rank substitute to correlation that could prove quite interesting in privacy settings, albeit no invariant to reparameterisations (with potential connections with Wasserstein distances), Sonia Petrone gave an overview on her substantial amount of work on empirical Bayes in Bayes (EBIB), making me wonder at natural ABC or EM ways to bypass the computation of the empirical Bayes hyperparameter computation (but also on measuring the overfitting degree of EBIB-ing), maybe exploiting the representation of the marginal as an average of predictive, and Debdeep Pati argued towards an interpretable and robust ML, using for estimation divergence a mixture of two KL’s that is remindful of GANs (as well as of Lasso and exponentially tilted empirical likelihood à la Chib & al.), implemented by a form of Bbootstrap. Plus enjoying a locally flavoured acronym, namely BETEL.
The second session I organised was centred on Bayesian computations, with both Sid Chib and Ritabrata Dutta (U Warwick) presenting a Bayesian modelling of difference-in-difference models in clinical trials and several works on (score based) generalized Bayes models with Shreya Roy (U Warwick), who further won a poster prize at IISA 2024. I terminated the session with a talk on our insufficient Gibbs sampler, which connected with some aspects of both Sid’s and Rito’s talks.

As in earlier editions of IISA I attended, the local organisation was most enjoyable, from supportive staff and students, relaxed atmosphere, easy commutes, heaps of great food, and unlimited chai! Plus meeting and listening to participants I had met in these earlier editions. Looking forward the 2026 edition!