Archive for RIKEN

Information Geometry, Privacy and Monte Carlo workshop, ISM, 6-7 July 2026

Posted in Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on July 8, 2026 by xi'an

Although some of the participants of the workshop left for ICML²⁶ or the 4th Bayesian Nonparametrics networking workshop, both taking place in Seoul this week, the following days of the workshop were as intense and captivating as the first two, with a return to MCMC “basics” but also more geometrical and maethematical aspects.

To wit, Radu Craiu talked on MCMC for DAG processes with revisiting the landmark paper of Geyer & Møller (1994) on replacing discrete time MCMC with a birth & death process and cutting on complexity by restricted set imposing some edges, set from a redetermined run. Galin Jones presented some (novel) Lower bounds on the rate of convergence for accept-reject-based Markov chains in Wasserstein and total variation distances, showing the massive dependence of the convergence rates on the scaling factors of the proposal, especially in relation with the data size n when considering posterior targets. James Flegal discussed Simultaneous confidence bands for (MC)MC simulations that aimed at returning a confidence band on marginal density estimates; it reminded me of our 2005 simultaneous coverage paper with Wilfrid Kendall and Jean-Michel Marin and got me wondering why not going full Bayes by adopting a GP prior modelling.

Michiko Okudo spoke about Applications of information geometry to Bayesian prediction and estimation in curved exponential families, returning to point estimation with a mention of Marchand & Strawderman (2025)! Marta Catalano presented results on Distances on random measures for Bayesian nonparametrics, involving random measures like Dirichlet processes, that was connected with Hugo Lavenant’s talk at ISBA, but more focussed on the mathematical aspects albeit algorithmic aspects were mentioned. With highly intuitive arguments (making the accronym WoW for Wasserstein on Wasserstein quite appropriate!).

Takemasa Miyoshi made a presentation of the Osaka Expo 2025 Weather [prediction] on Fugaku: Synergizing Big Data Assimilation and AIRIKEN, with impressive predictive abilities achieved using RIKEN super-computer (but no technical details). Björn Sprung exposed how they obtained Dimension-independent MCMC [convergence speed] on the sphere, using retroprojections of random walks outside the sphere (as in Frederica’s talk yesterday), which comes as a surprise given the deterioration of random walk performances with increasing dimensions.

Geoffrey Wolfer’s Characterization of Exponential Families of Lumpable Stochastic Matrices was a very mathematical talk set firmly in the Japanese probability school, going too fast with too many new definitions for my abilities (and attention span) but setting the scene for exponential families on stochastic matrices and being one of the rate cases I eve rsaw lumpability à la Kemeny & Snell (1983) mentionned! Daniel Paulin followed with Stochastic gradient Langevin dynamics: convergence and bias, via an UBU algorithm using splitting integrators that sound very much like the leapfrog for an HMC with unscented Langevin steps where the gradient is replaced with an unbiased estimator (connecting to the poster of Jack Jewson on Sunday, when he mentioned the opposition between pseudo-marginal MCMC, requiring an unbiased estimator of the target, and schemes using the log-target, for which unbiased estimators of the log can be used). Shahab Asoodeh concluded Monday with Recent Advances in Metropolis-Hastings Algorithms, actually developing multi-marginal coupling with freely coupling chains.

On the final morning, Weiming Feng showed results about a Faster mixing of the Jerrum-Sinclair chain, reminding me of the 1989 paper, with a Metropolis algorithm on graphs allowing for specific mixing time results with spectral gap and log-Sobolev inequalities (and a Poincáre typo!). Michael Choi produced convergence properties by Optimising two-block averaging kernels to speed up Markov chains, with a (rather formal) Gibbs sampler on orbits (in a finite state space) again connecting to Jerrum.

Yuga Iguchi discussed Diffusion models for high-dimensional clustered data: Intrinsic-dimension adaptivity via Bayesian classification, producing a rigorous characterisation of the phase transition property of their diffusion denoising probabilistic model when the target is a mixture with separation constraints on the components, phase transition meaning that eventually concentrating on a single cluster as the forward diffusion moves toward pure noise. (Although being fully awake, having mostly recovered from the longest jetlag period ever, I had trouble understanding the process per se.) Edric Tam discussed Fundamental Limits to Neural Monte Carlo by returning to standard variance reduction techniques like stratifying and antithetic-ying (!) and applying normalising flows on them. Victor Elvira concluded the meeting by Rethinking self-normalized importance sampling, with a fun interlude of Eric Veach’s Oscars joke, but I unfortunately had to miss the end to gather my bags and leave for the Alps! But Victor should be in Paris in the Fall and hopfefully giving a talk at mostly Monte Carlo!

This workshop was most efficiently supported by the Institute of Statistical Mathematics and its staff, including over the weekend days! On a personal foodie note, the coffee breaks featured the same unbelievable matcha cakes (“Chez Kobe”) as at ISBA²⁶, we enjoyed a terrific full tofu dinner at Umenohana Tachikawa shop and there were plenty French (or pseudo-French) bakeries in Tachekima, enough to find rye (raimugi) bread for breakfast!

Bayesian learning

Posted in Statistics with tags , , , , , , , , on May 4, 2023 by xi'an

“…many well-known learning-algorithms, such as those used in optimization, deep learning, and machine learning in general, can now be derived directly following the above scheme using a single algorithm”

The One World ABC webinar today was delivered by Emtiyaz Khan (RIKEN), about the Bayesian Learning Rule, following Khan and Rue 2021 arXival on Bayesian learning. (It had a great intro featuring a video of the speaker’s daughter learning about the purpose of a ukulele in her first year!) The paper argues about a Bayesian interpretation/version of gradient descent algorithms, starting with Zellner’s (1988, the year I first met him!) identity that the posterior is solution to

\min_q \mathbb E_q[\ell(\theta,x)] + KL(q||\pi)

when ℓ is the likelihood and π the prior. This identity can be generalised to an arbitrary loss function (also dependent on the data)  replacing the likelihood and considered for a posterior chosen within an exponential family just as variational Bayes. Ending up with a posterior adapted to this target (in the KL sense). The optimal hyperparameter or pseudo-hyperparameter of this approximation can be recovered by some gradient algorithm, recovering as well stochastic gradient and Newton’s methods. While constructing a prior out of a loss function would have pleased the late Herman Rubin, this is not the case, but rater an approach to deriving a generalised Bayes distribution within a parametric family, including mixtures of Gaussians. At some point in the talk, the uncertainty endemic to the Bayesian approach seeped back into the picture, but since most of the intuition came from machine learning, I was somewhat lost at the nature of this uncertainty.

 

 

the Bayesian learning rule [One World ABC’minar, 27 April]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , on April 24, 2023 by xi'an

The next One World ABC seminar is taking place (on-line, requiring pre-registration) on 27 April, 9:30am UK time, with Mohammad Emtiyaz Khan (RIKEN-AIP, Tokyo) speaking about the Bayesian learning rule:

We show that many machine-learning algorithms are specific instances of a single algorithm called the Bayesian learning rule. The rule, derived from Bayesian principles, yields a wide-range of algorithms from fields such as optimization, deep learning, and graphical models. This includes classical algorithms such as ridge regression, Newton’s method, and Kalman filter, as well as modern deep-learning algorithms such as stochastic-gradient descent, RMSprop, and Dropout. The key idea in deriving such algorithms is to approximate the posterior using candidate distributions estimated by using natural gradients. Different candidate distributions result in different algorithms and further approximations to natural gradients give rise to variants of those algorithms. Our work not only unifies, generalizes, and improves existing algorithms, but also helps us design new ones.

Concentration and robustness of discrepancy-based ABC [One World ABC ‘minar, 28 April]

Posted in Statistics, University life with tags , , , , , , , , , , , on April 15, 2022 by xi'an

Our next speaker at the One World ABC Seminar will be Pierre Alquier, who will talk about “Concentration and robustness of discrepancy-based ABC“, on Thursday April 28, at 9.30am UK time, with an abstract reported below.
Approximate Bayesian Computation (ABC) typically employs summary statistics to measure the discrepancy among the observed data and the synthetic data generated from each proposed value of the parameter of interest. However, finding good summary statistics (that are close to sufficiency) is non-trivial for most of the models for which ABC is needed. In this paper, we investigate the properties of ABC based on integral probability semi-metrics, including MMD and Wasserstein distances. We exhibit conditions ensuring the contraction of the approximate posterior. Moreover, we prove that MMD with an adequate kernel leads to very strong robustness properties.

approximate Bayesian inference [survey]

Posted in Statistics with tags , , , , , , , , , , , , , , , , , , on May 3, 2021 by xi'an

In connection with the special issue of Entropy I mentioned a while ago, Pierre Alquier (formerly of CREST) has written an introduction to the topic of approximate Bayesian inference that is worth advertising (and freely-available as well). Its reference list is particularly relevant. (The deadline for submissions is 21 June,)