Archive for boulevard périphérique

mostly Monte Carlo [last session of 2025]

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , on December 10, 2025 by xi'an

Rather hurriedly, here is the announcement for the last mostly MC seminar this year, to take place at 3-5pm this very Friday, Dec 12, 2025. It will take place in Salle 03, PariSanté Campus


3pm Least squares variational inference

Yvann Le Fay CREST, ENSAE

Variational inference seeks the best approximation of a target distribution within a chosen family, where “best” means minimizing Kullback-Leibler divergence. When the approximation family is exponential, the optimal approximation satisfies a fixed-point equation. We introduce LSVI (Least Squares Variational Inference), a gradient-free, Monte Carlo-based scheme for the fixed-point recursion, where each iteration boils down to performing ordinary least squares regression on tempered log-target evaluations under the variational approximation. We show that LSVI is equivalent to biased stochastic natural gradient descent and use this to derive convergence rates with respect to the numbers of samples and iterations. When the approximation family is Gaussian, LSVI involves inverting the Fisher information matrix, whose size grows quadratically with dimension d. We exploit the regression formulation to eliminate the need for this inversion, yielding O(d³) complexity in the full-covariance case and O(d) in the mean-field case. Finally, we numerically demonstrate LSVI’s performance on various tasks, including logistic regression, discrete variable selection, and Bayesian synthetic likelihood, showing competitive results with state-of-the-art methods, even when gradients are unavailable.

4pm Beyond the Unified Skew-Normal: Extended Models for Bayesian Classification

Paolo Onorati CEREMADE, Université Paris Dauphine – PSL

Binary classification models typically lose the conjugacy and computational simplicity enjoyed by Gaussian models. While the Unified Skew-Normal (SUN) family has recently been shown to be conjugated under the probit model, two new developments are presented that extend this idea to a broader class of link functions, including both logit and probit. In the parametric setting, the Perturbed Unified Skew-Normal (pSUN) distribution is introduced; it is conjugate to any binary regression model whose link admits a scale-mixture representation of Gaussian random variables, enabling tractable posterior summaries, efficient sampling schemes, and strong performance in high-dimensional covariate settings. The discussion then moves to the nonparametric domain, where the Quasi SUN family and the associated stochastic process provide conjugacy for nonparametric logit and probit models while preserving key closure properties. A stochastic representation of this process yields practical computational improvements over existing Gaussian-based approximations. Together, these SUN-type extensions offer promising tools for Bayesian classification with accurate posterior inference.

La Grande Course RATP du Grand Paris [1:29:13, 482/4415, 2/81 M5M]

Posted in pictures, Running with tags , , , , , , , , , , , , , , on April 3, 2025 by xi'an

 

new campus

Posted in pictures, Running, Travel, University life with tags , , , , , , , , , , , , on September 4, 2022 by xi'an

While I am keeping my office at Porte Dauphine, undergoing major renovations (of the 1955 NATO building!), I am now spending most of my time in a more modern campus, called PariSanté, located at Porte de Versailles, with medical research teams and startups. This is where our master MASH will be located. The place is very luminous and despite the close proximity with the Paris beltway (le périf’), quiet (and much quieter than Paris Dauphine). It is also an ecological absurdity, with a huge sunroof that could not be shaded during the heat waves, plastic trees, self-induced lights, and compulsory lifts. On the memory lane, it is a trip back 35 years ago, as it sits across the road from the Balard military compound where I spent most of my military service in 1987 (working on my PhD in a research department).  And it is conveniently located half-way between home and Paris Dauphine, although not skipping the tough hill of Porte de Versailles on the way back..!

to be demolished!

Posted in Books, pictures, University life with tags , , , , , , , , , on May 8, 2022 by xi'an

Tractable Fully Bayesian inference via convex optimization and optimal transport theory

Posted in Books, Statistics, University life with tags , , , , , , , , on October 6, 2015 by xi'an

IMG_0294“Recently, El Moselhy et al. proposed a method to construct a map that pushed forward the prior measure to the posterior measure, casting Bayesian inference as an optimal transport problem. Namely, the constructed map transforms a random variable distributed according to the prior into another random variable distributed according to the posterior. This approach is conceptually different from previous methods, including sampling and approximation methods.”

Yesterday, Kim et al. arXived a paper with the above title, linking transport theory with Bayesian inference. Rather strangely, they motivate the transport theory with Galton’s quincunx, when the apparatus is a discrete version of the inverse cdf transform… Of course, in higher dimensions, there is no longer a straightforward transform and the paper shows (or recalls) that there exists a unique solution with positive Jacobian for log-concave posteriors. For instance, log-concave priors and likelihoods. This solution remains however a virtual notion in practice and an approximation is constructed via a (finite) functional polynomial basis. And minimising an empirical version of the Kullback-Leibler distance.

I am somewhat uncertain as to how and why apply such a transform to simulations from the prior (which thus has to be proper). Producing simulations from the posterior certainly is a traditional way to approximate Bayesian inference and this is thus one approach to this simulation. However, the discussion of the advantage of this approach over, say, MCMC, is quite limited. There is no comparison with alternative simulation or non-simulation methods and the computing time for the transport function derivation. And on the impact of the dimension of the parameter space on the computing time. In connection with recent discussions on probabilistic numerics and super-optimal convergence rates, Given that it relies on simulations, I doubt optimal transport can do better than O(√n) rates. One side remark about deriving posterior credible regions from (HPD)  prior credible regions: there is no reason the resulting region is optimal in volume (HPD) given that the transform is non-linear.