Archive for stochastic optimisation

BayesComp 2025.4

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on June 21, 2025 by xi'an

The third and final day of the (main) conference started tih Emtiyaz Khan’s plenary talk on adaptive Bayesian intelligence. Or, imho, [adaptive [Bayesian]] intelligence, with the brackets indicating redundancy since intelligence need include adaptivity and [intelligent] adaptivity need proceed in a Bayesian way! Focussing first on the Bayesian learning rule via variational Bayes (with a stress on Kingma’s 1994 Adam optimisation algorithm, the “most cited paper” [in machine learning]) where learning boils down to gradient steps (due to the exponential family structure), themselves versions of Taylor (or Laplace) approximations). With an interesting vision of Bayesian updating as accounting for prediction mismatch. (I missed the connection Roberta in IMDb appearing in one slide!)

 The following session offered no dilemma [sorry, Alex, Axel, Chris, Robert, Sumeet, Victor!] since it included the federated learning session I organised, with Louis Asslet, Conor Hassan, and Jean-Michel Marin as speakers. Louis’ talk was on confidential [homomorphic] accept-reject algorithms to learn from other sources, while preserving (differential?) privacy, part of which came during Les Houches workshops I organised this Spring and the one before. Exploiting the additive features of log-likelihoods and exponential variates and adopting a testing perspective on privacy. Conor motivated his model with the Australian cancer atlas project Kerrie Mengersen and others have been developing over the years. The federated approach relies on variational approximations that return the same answer as an exact resolution, but more efficiently. (From a privacy perspective, I wonder at the impact of variational approximations on protecting the data, which boils down to a choice of (sufficient) statistics for the exponential families behind those approximations.) For more complicated models incorporating spatial dependence prohibits full Bayesian inference, unfortunately. Jean-Michel commented on the richness of methods for simulation-based inference, incl. model choice. His focus was on using sequential neural likelihood estimation and sequential importance sampling to approximate evidence. As in the Read Paper of Del Moral et al. (2006). Mentioning a neural version of the harmonic mean estimator by Spurio Mancini et al.  (2023)! I wondered at the degree of (Rao-Blackwell) recycling involved in the computation, Jean-Michel’s answer being that AMIS is soon coming [in a theatre near you!].

The afternoon sessions did offer any reprieve in the choice of topic! I first went to Approximate Methods for Accelerated Sampling, with Rong Tang evaluating the informativeness of summary statistics through a divergence evaluation. Using autoencoders to replace the intractable posterior, with sliced minimal model discrepancy (MMD) and (pseudo?) score matching loss for divergences (reminding me of indirect inference and synthetic likelihood). Yun Yang discussed a variational proposal to estimate the number of components in a mixture model. Surprising given the multimodal structure of mixture posteriors. And the overall irregularity of (evil!) mixture models. But I could not figure out from the talk the form of the approximation.

On the food scene, tasted a nice and spicy Peranakan rice vermicelli dish called Mee Siam yesterday in a campus restaurant, which sustained me fore the rest of the day, including the ABC s/webinar. And another spicy hot pot today at NUS, to catch up on veggies, while missing the chili crab local specialty on that trip.

optimal importance sampling for stochastic optimisation

Posted in Books, Statistics, University life with tags , , , , , , on June 6, 2025 by xi'an

A recent arXival by Liviu Aolaritei, Bart Van Parys, Henry Lam, and Michael Jordan (a co-PI in our ERC Synergy Ocean project) discusses optimal importance sampling schemes for stochastic optimisation, processed by an iterative Robbins-Munro algorithm improvement (with the Polyak-Ruppert improvement).

“Despite its popularity, IS is often described as a `double-edged sword.’ Its performance depends critically on the choice of the proposal distribution, which is typically sensitive to the underlying model”

I had never thought of optimising the importance function in this context, even after the seminaire of Tom Guédon mentioned in a recent ‘og.  The paper considers the optimisation of f(θ)=E[F(θ,X)] whose expectation is under a certain distribution P, through

“an iterative gradient-based algorithm that jointly updates the decision variable and the IS distribution without requiring time-scale separation between the two. Our method achieves the lowest possible asymptotic variance and guarantees global convergence under convexity of the objective and mild assumptions on the IS distribution family. Furthermore, we show that these properties are preserved under linear constraints by incorporating a recent variant of Nesterov’s dual averaging method”

with the Robbins-Munro algorithm updating the value of both θ and the parameter of the importance function. One specific difficulty is that the ideal importance function depends on the argument of the optimisation problem, obviously unknown, a “curse of circularity” I had not met previously. The linear constraint creates another kind of difficulty in order to derive the active constraint set while requiring an adapted (eg, absolutely continuous) importance function. This leads the authors to introduce a different (secondary) importance distribution on X, opening a Pandora box of infinite tuning that they opt to terminate by fixing an importance function at some finite stage. I am however uncertain as to how the combinatoric difficulty of exploring all active constraint sets at each iteration is handled. The paper being mostly theoretical, there is no illustration therein. Nor a computational cost evaluation.

Stochastic Numerics and Statistical Learning

Posted in Statistics, Travel, University life with tags , , , , on May 25, 2023 by xi'an

This week, I am taking part in the Stochastic Numerics and Statistical Learning workshop, in KAUST, which should prove an interesting experience!

far south

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , , , on February 23, 2022 by xi'an

ICM 2018

Posted in pictures, Statistics, Travel, University life with tags , , , , , , , on August 4, 2018 by xi'an

While I am not following the International Congress of Mathematicians which just started in Rio, and even less attending, I noticed an entry on their webpage on my friend and colleague Maria Esteban which I would have liked to repost verbatim but cannot figure how. (ICM 2018 also features a plenary lecture by Michael Jordan on gradient based optimisation [which was also Michael’s topic at ISBA 2018] and another one by Sanjeev Arora on the maths deep learning, two talks broadly related with statistics, which is presumably a première at this highly selective maths conference!)