Archive for particle filters

Information Geometry, Privacy and Monte Carlo workshop, ISM, 6-7 July 2026

Posted in Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on July 8, 2026 by xi'an

Although some of the participants of the workshop left for ICML²⁶ or the 4th Bayesian Nonparametrics networking workshop, both taking place in Seoul this week, the following days of the workshop were as intense and captivating as the first two, with a return to MCMC “basics” but also more geometrical and maethematical aspects.

To wit, Radu Craiu talked on MCMC for DAG processes with revisiting the landmark paper of Geyer & Møller (1994) on replacing discrete time MCMC with a birth & death process and cutting on complexity by restricted set imposing some edges, set from a redetermined run. Galin Jones presented some (novel) Lower bounds on the rate of convergence for accept-reject-based Markov chains in Wasserstein and total variation distances, showing the massive dependence of the convergence rates on the scaling factors of the proposal, especially in relation with the data size n when considering posterior targets. James Flegal discussed Simultaneous confidence bands for (MC)MC simulations that aimed at returning a confidence band on marginal density estimates; it reminded me of our 2005 simultaneous coverage paper with Wilfrid Kendall and Jean-Michel Marin and got me wondering why not going full Bayes by adopting a GP prior modelling.

Michiko Okudo spoke about Applications of information geometry to Bayesian prediction and estimation in curved exponential families, returning to point estimation with a mention of Marchand & Strawderman (2025)! Marta Catalano presented results on Distances on random measures for Bayesian nonparametrics, involving random measures like Dirichlet processes, that was connected with Hugo Lavenant’s talk at ISBA, but more focussed on the mathematical aspects albeit algorithmic aspects were mentioned. With highly intuitive arguments (making the accronym WoW for Wasserstein on Wasserstein quite appropriate!).

Takemasa Miyoshi made a presentation of the Osaka Expo 2025 Weather [prediction] on Fugaku: Synergizing Big Data Assimilation and AIRIKEN, with impressive predictive abilities achieved using RIKEN super-computer (but no technical details). Björn Sprung exposed how they obtained Dimension-independent MCMC [convergence speed] on the sphere, using retroprojections of random walks outside the sphere (as in Frederica’s talk yesterday), which comes as a surprise given the deterioration of random walk performances with increasing dimensions.

Geoffrey Wolfer’s Characterization of Exponential Families of Lumpable Stochastic Matrices was a very mathematical talk set firmly in the Japanese probability school, going too fast with too many new definitions for my abilities (and attention span) but setting the scene for exponential families on stochastic matrices and being one of the rate cases I eve rsaw lumpability à la Kemeny & Snell (1983) mentionned! Daniel Paulin followed with Stochastic gradient Langevin dynamics: convergence and bias, via an UBU algorithm using splitting integrators that sound very much like the leapfrog for an HMC with unscented Langevin steps where the gradient is replaced with an unbiased estimator (connecting to the poster of Jack Jewson on Sunday, when he mentioned the opposition between pseudo-marginal MCMC, requiring an unbiased estimator of the target, and schemes using the log-target, for which unbiased estimators of the log can be used). Shahab Asoodeh concluded Monday with Recent Advances in Metropolis-Hastings Algorithms, actually developing multi-marginal coupling with freely coupling chains.

On the final morning, Weiming Feng showed results about a Faster mixing of the Jerrum-Sinclair chain, reminding me of the 1989 paper, with a Metropolis algorithm on graphs allowing for specific mixing time results with spectral gap and log-Sobolev inequalities (and a Poincáre typo!). Michael Choi produced convergence properties by Optimising two-block averaging kernels to speed up Markov chains, with a (rather formal) Gibbs sampler on orbits (in a finite state space) again connecting to Jerrum.

Yuga Iguchi discussed Diffusion models for high-dimensional clustered data: Intrinsic-dimension adaptivity via Bayesian classification, producing a rigorous characterisation of the phase transition property of their diffusion denoising probabilistic model when the target is a mixture with separation constraints on the components, phase transition meaning that eventually concentrating on a single cluster as the forward diffusion moves toward pure noise. (Although being fully awake, having mostly recovered from the longest jetlag period ever, I had trouble understanding the process per se.) Edric Tam discussed Fundamental Limits to Neural Monte Carlo by returning to standard variance reduction techniques like stratifying and antithetic-ying (!) and applying normalising flows on them. Victor Elvira concluded the meeting by Rethinking self-normalized importance sampling, with a fun interlude of Eric Veach’s Oscars joke, but I unfortunately had to miss the end to gather my bags and leave for the Alps! But Victor should be in Paris in the Fall and hopfefully giving a talk at mostly Monte Carlo!

This workshop was most efficiently supported by the Institute of Statistical Mathematics and its staff, including over the weekend days! On a personal foodie note, the coffee breaks featured the same unbelievable matcha cakes (“Chez Kobe”) as at ISBA²⁶, we enjoyed a terrific full tofu dinner at Umenohana Tachikawa shop and there were plenty French (or pseudo-French) bakeries in Tachekima, enough to find rye (raimugi) bread for breakfast!

Information Geometry, Privacy and Monte Carlo workshop, ISM, 4-5 July 2026

Posted in Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , on July 6, 2026 by xi'an

After the (exciting) variety and spread of ISBA²⁶, here we are at much more focussed (and single-track), if equally exciting, workshop at the ISM. (With many participants from ISBA²⁶.)

On Saturday afternoon, Ajay Jasra talked about Particle filtering for state-space models with low, degenerate noise, with specific measure issues I did not really get, since the manifold attached to the noise was known, but the projected density may prove a challenge. Manifolds were also central to Kenji Fukumizu’s talk on Learning manifold structure and density with score-based models learning scores as projectors to the manifold, although it was unclear to me how this was possible when the manifold is unknown. Christophe Andrieu presented Geometry informed selection in multiple proposal MCMC which stems from an early multi-proposal (1998) proposal by Radford Neal and uses a multivariate ranking procedure to quantify a measure of surprise for the current  Markov chain value within the proposed ones. The crux for the efficiency of the approach may be in the choice of this ranking procedure. And Maria De Iorio talked about Efficient MCMC via similarity-driven proposals for discrete support targets, with similarities with ABC.

Completed with a human sized poster session where I reconnected with Spanish friends I had not seen for ages (by missing OBayes meetings).

On (pleasantly rainy) Sunday morning, Federica Milinanni detailled her Rapid mixing of stereographic MCMC for heavy-tailed sampling, essentially the same content as in Nagoya last FRiday, with a novel sub-Cauchy projection supposed to explore heavy tails better: while the regular stereographic projection turns the t-distribution with d degrees of freedom into a Uniform on the hypersphere, a sub-Cauchy projection turns the Cauchy into this uniform. In the privacy session I organised, Hongsheng Dai, member of our ERC Synergy project, presented an Online federated learning framework for classification, using DP as a criterion and achieving by adding noise to the loss function at each occurrence of the data production. Surprisingly increasing with the number of occurrences, not so much since the objective function keeps calling

Stefano Favaro described his Bayesian nonparametric privacy-preserving synthetic data generation method (for discrete data) that connects privacy protection and information preservation. (Incidentally I was unaware of the σ parameter of the Pitman-Yor process, which allows for a finite support when σ<0, but I cannot fathom the appeal of this extension, given the complete lack of connection between the positive and negative cases.) Surprisingly, non-parametric prediction does worse in terms of privacy, if not so surprising with discrete data since the predictive actually put weight on every datapoint. Resorting to  mechanism informativity by Wasserman and Zhou (2010) (with a surprise mention of my friend Arnaud Guilin!). And Joshua Bon gave his Persuasive Privacy talk of last Tuesday  (to be re-repeated two days later at ICML²⁶ in Seoul!). Except for changing the audience game from croissants to sumo wrestlers! (What will it be in Seoul!?) And adding much more details on the foundational elements of persuasive privacy.

The poster session was similarly enjoyable, even though I did not manage to get through all posters.

BayesComp 2025.3

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on June 20, 2025 by xi'an

The second day of the conference started with a cooler and less humid weather (although this did not last!), although my brain felt a wee bit foggy from a lack of sleep (and I almost crashed while running on the hotel treadmill, at 14.5km/h!), and the plenary talk of my friend of many years Sylvia Früwirth-Schnatter on horseshoe priors and time-varying time series (à la West). With a nice closed-form representation involving hypergeometric functions of the second kind (my favourite!), with the addition of a triple-Gamma prior. Sylvia stressed on the enormous impact of the prior choice on change-point detection, which was already the point in the original horseshoe paper (as opposed to George’s Lasso prior). Without incorporating any specific modelling on potential change-point, fair enough given that the parameter is moving with time, unhindered. Her MCMC choices involved discrete parameters with Negative Binomial and Poisson parameters, allowing for partially integrated or collapsed solutions. Possibly further improved by Swendsen-Wang steps.

I then attended the (advanced) Langevin session after agonising upon my choice for a wealth of options! Sam Power presented a talk linking simulation with optimisation targets, over measure spaces. With Wasserstein gradient flow algorithms that resemble Langevin algorithms once discretised by a particle system. (A natural resolution producing a somewhat unnatural form of measure estimator since made of Dirac masses, from which very little can be learned.) Then [my Warwick colleague & coauthor] Any Wang on underdamped Langevin diffusions. when Poincaré‘s inequality fails, but convergence (in total variation) still occurs. Followed by Peter Whalley on splitting methods (where random hypergeometric subsampling dominates Robbins-Monro) and stochastic gradient algorithms, in a connected (to the previous talks) way since involving underdamped aspects. (With a personal discovery of Polyak’s heavy ball method.)

The afternoon session saw me facing a terrible dilemma with three close friends talking at the same time! Eventually opting for PDMPs, over simulation-based inference and recalibration for approximate Bayesian methods. Kengo Kamatani gave a general introduction to PDMPs, before explaining the automated implementation he considered with Charly Andral (during Charly’s visit to ISM, Tokyo, two summers ago). Towards accelerating the generation of the jump time. Then Luke Hardcastle applied PDMPs for survival prediction, using spike & slab priors and sticky PDMPs. And Jere Koskela (formerly Warwick) extended zig-zag sampling to discrete settings (incl. Kingman’s coalescent.)

The (rather long) day was not over yet since we had planned an extra on-site OWABI seminar & webinar with two participants in the conference, Filippo Pagani (Warwick and OCEAN postdoc) using fusion for federated learning, with a trapezoidal approximation, and Maurizio Filippone on GANs as hidden perfect ABC model selection, a GAN providing an automatic density estimator… With astounding Gemini-generated cartoons! Videos are soon to be available. A big congrats to the speakers who managed to convey their ideas and results despite the late hour! (On the extra-academic side, I was invited last night to a genuine Szechuan dinner in Chinatown, with a large array of spicy dishes if not that spicy!, and a rare opportunity to taste abalone. And bullfrogs. Quite a treat! And a good reason to skip dinner altogether!)

SMC 22 coming soon!

Posted in Statistics with tags , , , , , , , , , on February 7, 2022 by xi'an

The 5th Workshop on Sequential Monte Carlo Methods (SMC 2022) will take place in Madrid on 4-6 May 2022. More precisely on the Leganés campus of Universidad Carlos III de Madrid. Registrations are now open, with very modest registration fees and the list of invited speakers is available on the webpage of the workshop. (The SMC 2020 workshop was cancelled due to the COVID-19 pandemic. An earlier workshop took place at CREST in 2015.)

EM degeneracy

Posted in pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , on June 16, 2021 by xi'an

At the MHC 2021 conference today (to which I biked to attend for real!, first time since BayesComp!) I listened to Christophe Biernacki exposing the dangers of EM applied to mixtures in the presence of missing data, namely that the algorithm has a rising probability to reach a degenerate solution, namely a single observation component. Rising in the proportion of missing data. This is not hugely surprising as there is a real (global) mode at this solution. If one observation components are prohibited, they should not be accepted in the EM update. Just as in Bayesian analyses with improper priors, the likelihood should bar single or double  observations components… Which of course makes EM harder to implement. Or not?! MCEM, SEM and Gibbs are obviously straightforward to modify in this case.

Judith Rousseau also gave a fascinating talk on the properties of non-parametric mixtures, from a surprisingly light set of conditions for identifiability to posterior consistency . With an interesting use of several priors simultaneously that is a particular case of the cut models. Namely a correct joint distribution that cannot be a posterior, although this does not impact simulation issues. And a nice trick turning a hidden Markov chain into a fully finite hidden Markov chain as it is sufficient to recover a Bernstein von Mises asymptotic. If inefficient. Sylvain LeCorff presented a pseudo-marginal sequential sampler for smoothing, when the transition densities are replaced by unbiased estimators. With connection with approximate Bayesian computation smoothing. This proves harder than I first imagined because of the backward-sampling operations…