Archive for Bayesian classification

robust simulation-based inference

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , on March 7, 2026 by xi'an

This new arXival by Lorenzo Tomaselli, Valérie Ventura, and Larry Wasserman (from CMU) considers simulation-based inference under model misspecification (as we did for ABC in our 2020 Series B paper). Which is almost always the case. In the paper, SBI is defined as producing N parameters and N samples from the prior and the corresponding sampling distribution, respectively, and then doubling the resulting samples by permuting at random the parameters θ. This means that the second half is distributed from the product of the prior and of the marginal, hence that the classification odds ratio is equal to the likelihood, hence providing an estimation method (andlikelihood trick) à la Geyer. From this estimate, an ABC p-value can be derived, but it is incorrect as such when the model is misspecified. Hence the use of the Hellinger discrepancy, the power divergence and the kernel distance (or MMD) as alternatives to the misspecified MLE.

The paper then expands on approximating density ratios by virtue of a reproducing kernel Hilbert space, using a Gaussian kernel. (With a nice remark on requiring only one single ratio estimator for all values of θ, albeit in the joint space.) And focus on a studentized MMD estimator (à la e-value) to build a confidence set that remains valid under model misspecification. And without regularity assumptions.

Another approach is further explored, based on exponential tilting—of which I am not a great fan, from being highly dependent on the choice of the pseudo-sufficient statistic to require an intractable normalising constant, to requiring an extra optimization, even though I appreciate the mathematical appeal of the construct. Which seems to require a sample simulation for each value of θ at the learning stage, albeit relying on the same likelihood trick. The appropriateness of the tilting can be tested by a goodness of fit test tailored for the SBI structure, which sounds rather greedy in the required simulations. 

Besides the g-and-k distribution example (which, as pointed out several times on the ‘Og, is not intractable, strictly speaking!), the paper studies a mixture example, despite Larry dubbing them as evil as tequila a long while ago! (The paper also offers a section called accoutrements, which is my first encounter with this use of the term, usually found in medieval contexts!)

Note that Larry will present the paper at the OWABI webinar next 25 March!

mostly Monte Carlo [last session of 2025]

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , on December 10, 2025 by xi'an

Rather hurriedly, here is the announcement for the last mostly MC seminar this year, to take place at 3-5pm this very Friday, Dec 12, 2025. It will take place in Salle 03, PariSanté Campus


3pm Least squares variational inference

Yvann Le Fay CREST, ENSAE

Variational inference seeks the best approximation of a target distribution within a chosen family, where “best” means minimizing Kullback-Leibler divergence. When the approximation family is exponential, the optimal approximation satisfies a fixed-point equation. We introduce LSVI (Least Squares Variational Inference), a gradient-free, Monte Carlo-based scheme for the fixed-point recursion, where each iteration boils down to performing ordinary least squares regression on tempered log-target evaluations under the variational approximation. We show that LSVI is equivalent to biased stochastic natural gradient descent and use this to derive convergence rates with respect to the numbers of samples and iterations. When the approximation family is Gaussian, LSVI involves inverting the Fisher information matrix, whose size grows quadratically with dimension d. We exploit the regression formulation to eliminate the need for this inversion, yielding O(d³) complexity in the full-covariance case and O(d) in the mean-field case. Finally, we numerically demonstrate LSVI’s performance on various tasks, including logistic regression, discrete variable selection, and Bayesian synthetic likelihood, showing competitive results with state-of-the-art methods, even when gradients are unavailable.

4pm Beyond the Unified Skew-Normal: Extended Models for Bayesian Classification

Paolo Onorati CEREMADE, Université Paris Dauphine – PSL

Binary classification models typically lose the conjugacy and computational simplicity enjoyed by Gaussian models. While the Unified Skew-Normal (SUN) family has recently been shown to be conjugated under the probit model, two new developments are presented that extend this idea to a broader class of link functions, including both logit and probit. In the parametric setting, the Perturbed Unified Skew-Normal (pSUN) distribution is introduced; it is conjugate to any binary regression model whose link admits a scale-mixture representation of Gaussian random variables, enabling tractable posterior summaries, efficient sampling schemes, and strong performance in high-dimensional covariate settings. The discussion then moves to the nonparametric domain, where the Quasi SUN family and the associated stochastic process provide conjugacy for nonparametric logit and probit models while preserving key closure properties. A stochastic representation of this process yields practical computational improvements over existing Gaussian-based approximations. Together, these SUN-type extensions offer promising tools for Bayesian classification with accurate posterior inference.