Archive for SMC

special issue of Statistica Sinica on sequential Monte Carlo

Posted in Books, Statistics, University life with tags , , , , , , , on May 24, 2024 by xi'an

6th Workshop on Sequential Monte Carlo Methods

Posted in Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , on May 16, 2024 by xi'an

Very glad to be back to an SMC workshop as it has been nine years since my attending SMC 2015 in Malakoff! The more for the workshop taking place in Edinburgh and at the Bayes Centre. It is one of these places where I feel somewhat returning to familiar grounds with accumulated memories. Like my last visit there when I had a tea with Mike Titterington…

The overall pace of the workshop was quite nice, with long breaks for informal discussions (and time for ‘oggin’!) and interesting poster late afternoons, helped by the small number of them at each instance, incl. one on reversible jump HMC. Here are a few scribbled entries about some talks along the first two days.

After my opening talk (!), Joaquín Míguez talked about the impact of a sequential (Euler-Marayama) discretisation scheme for stochastic differential equations on Bayesian filtering with control of the approximation effect. Axel Finke (in a joint work with Adrien Corenflos, now an ERC Ocean postdoc in Warwick) built a sequence of particle filter algorithms targeting good performances (high expected jumping distance) against both large dimensions and high time horizon, exploiting gradient shift MALA-like, as well as prior impact, with the conclusion that their jack-of-all-trades solutions, Particle­-MALA and Particle­-mGRAD, enjoyed this resistance in nearly normal models. Interesting reminder of the auxiliary particle trick and good insights on using the smoothing target, even when accounting for the computing time, but too many versions for a single talk without checking against the preprint.

The SMC sampler-like algorithm involves propagating N “seed” particles z(i), with a mutation mechanism consisting of the generation of N integrator snippets 𝗓:=(z,ψ⁢(z),ψ²⁢(z),…) started at every seed particle z(i), resulting in N×(T+1) particles which are then whittled down to a set of N seed particles using a standard resampling scheme. Andrieu et al., 2024

Christophe Andrieu talked about Monte Carlo sampling with integrator snippets, starting with recycling solutions for the leapfrog integrator HMC and unfolding Hamiltonians for moving more easily. With snippets representing discretised paths along the level sets being used as particles, picking zero, one, or more particles along each path, since importance weights are connection with multinomial HMC

This relatively small algorithmic modification of the conditional particle filter, which we call the conditional backward sampling particle filter has a dramatically improved performance over the conditional particle filter. Karjalainen et al., 2024

Anthony Lee looked at mixing times for backward sampling SMC (CBPF/ancestor sampling) cf Lee et al. (2020), where the backward step consists in computing the weight of a randomly drawn backward or ancestral history. Improving on earlier results to reach mixing time O(log T) and complexity O(T log T) (with T the time horizon). Thanks to maximal coupling and boundedness assumptions on the prior and likelihood functions.

Neil Chada presented a work on Bayesian multilevel Monte Carlo on deep networks. À la Giles, with a telescoping identity. Always puzzling to envision a prior on all parameters of a neural network. Achieving a computational cost inverse to the order of the MSE, at best. With a useful reminder that pushing the size of the NN to infinity results in a (poor) Gaussian process prior (Sell et al., 2023).

On my first evening, I stopped with a friend in my favourite Blonde [restaurant], as in almost every other visit to Edinburgh, enjoyable as always, but I also found the huge offer of Asian minimarkets in the area too tempting to resist, between Indian, Korean, and Chinese products. (Although with a disappointing hojicha!). As I could not reach any new Munro by train or bus within a reasonable time range I resorted to the nearer Pentland Hills, with a stop by Rosslyn Chapel (mostly of Da Vinci Code fame!, if classic enough). And some delays in finding a bus getting there (misled by google map!) and a trail (misled by my poor map reading skills) up the actual hills. The mist did not help either.

connection between tempering & entropic mirror descent

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , on April 30, 2024 by xi'an

The next One World ABC webinar is this  Thursday,  the 2nd May, at 9am UK time, with Francesca Crucinio (King’s College London, formerly CREST and even more formerly Warwick) presenting

“A connection between Tempering and Entropic Mirror Descent”.

a joint work with Nicolas Chopin and Anna Korba (both from CREST) whose abstract follows:

This work explores the connections between tempering (for Sequential Monte Carlo; SMC) and entropic mirror descent to sample from a target probability distribution whose unnormalized density is known. We establish that tempering SMC corresponds to entropic mirror descent applied to the reverse Kullback-Leibler (KL) divergence and obtain convergence rates for the tempering iterates. Our result motivates the tempering iterates from an optimization point of view, showing that tempering can be seen as a descent scheme of the KL divergence with respect to the Fisher-Rao geometry, in contrast to Langevin dynamics that perform descent of the KL with respect to the Wasserstein-2 geometry. We exploit the connection between tempering and mirror descent iterates to justify common practices in SMC and derive adaptive tempering rules that improve over other alternative benchmarks in the literature.

simulation as optimization [by kernel gradient descent]

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , on April 13, 2024 by xi'an

Yesterday, which proved an unseasonal bright, warm, day, I biked (with a new wheel!) to the east of Paris—in the Gare de Lyon district where I lived for three years in the 1980’s—to attend a Mokaplan seminar at INRIA Paris, where Anna Korba (CREST, to which I am also affiliated) talked about sampling through optimization of discrepancies.
This proved a most formative hour as I had not seen this perspective earlier (or possibly had forgotten about it). Except through some of the talks at the Flatiron Institute on Transport, Diffusions, and Sampling last year. Incl. Marilou Gabrié’s and Arnaud Doucet’s.
The concept behind remains attractive to me, at least conceptually, since it consists in approximating the target distribution, known up to a constant (a setting I have always felt standard simulation techniques was not exploiting to the maximum) or through a sample (a setting less convincing since the sample from the target is already there), via a sequence of (particle approximated) distributions when using the discrepancy between the current distribution and the target or gradient thereof to move the particles. (With no randomness in the Kernel Stein Discrepancy Descent algorithm.)
Ana Korba spoke about practically running the algorithm, as well as about convexity properties and some convergence results (with mixed performances for the Stein kernel, as opposed to SVGD). I remain definitely curious about the method like the (ergodic) distribution of the endpoints, the actual gain against an MCMC sample when accounting for computing time, the improvement above the empirical distribution when using a sample from π and its ecdf as the substitute for π, and the meaning of an error estimation in this context.

“exponential convergence (of the KL) for the SVGD gradient flow does not hold whenever π has exponential tails and the derivatives of ∇ log π and k grow at most at a polynomial rate”

futuristic statistical science [editorial]

Posted in Books, Kids, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , on January 13, 2024 by xi'an

This special issue of Statistical Science is devoted to the future of Bayesian computational statistics, from several perspectives. It involves a large group of researchers who contributed to collective articles, bringing their own perspectives and research interests into these surveys. Somewhat paradoxically, it starts with the past—and a conference on a Gold Coast beach. Martin, Frazier, and Robert first submitted a survey on the history of Bayesian computation, written after Gael Martin delivered a plenary lecture at Bayes on the Beach, a conference held in November 2017 in Surfers Paradise, Gold Coast, Queensland, and organised by Bayesian Research and Applications Group (BRAG), the Bayesian research group headed by Kerrie Mengersen at the Queensland University of Technology (QUT). Following a first round of reviews, this paper got split into two separate articles, Computing Bayes: From Then ‘Til Now , retracing some of the history of Bayesian computation, and Approximating Bayes in the 21st Century, which is both a survey and a prospective on the directions and trends of approximate Bayesian approaches (and not solely ABC). At this point, Sonia Petrone, editor of Statistical Science, suggested we had a special issue on the whole issue of trends of interest and promise for Bayesian computational statistics. Joining forces, after some delays and failures to convince others to engage, or to produce multilevel papers with distinct vignettes, we eventually put together an additional four papers, where lead authors gathered further authors to produce this diverse picture of some incoming advances in the field. We have deliberated avoided topics which have excellent recent reviews— such as Stein’s method, sequential Monte Carlo, piecewise deterministic Markov processes— and topics which are still in their infancy, such as the relationship of Bayesian approaches to large language models (LLMs) and foundation models.

Within this issue, Past, Present, and Future of Software for Bayesian Inference from Erik Štrumbelj & al covers the state of the art in the most popular Bayesian software, reminding us of the massive impact BUGS has had on the adoption of Bayesian tools since its early introduction in the early 1990s (which I remember discovering at the Fourth Valencia meeting on Bayesian statistics in April 1991). With an interesting distinction between first and second generations, and a light foray of the potential third generation, maybe missing the role of LLMs in coding that are already impacting the approach to computing and the less immediate revolution brought by quantum computing. Winter & al.’s The Future of Bayesian Computation [TITLE TO CHANCE] is making a link with machine learning techniques, without looking at the scariest issue of how Bayesian inference can survive in a machine learning world! While it produces an additional foray into the blurry division between proper sampling (à la MCMC) and approximations, additional to the historical Martin et al. (2024), it articulates these aspects within a (deep) machine learning perspective, emphasizing the role of summaries produced by generative models exploiting the power of neural network computation/optimization. And the pivotal reliance on variational Bayes, which is the most active common denominator with machine learning. With further entries on major issues like distributed computing, opening on the important aspect of data protection and guaranteed  privacy. We particularly like the clinical presentation of this paper with attention to automation and limitations. Normalizing flows actually link this paper with Heng, Bortoli and Doucet’s coverage of the Schrödinger bridge, which is a more focussed coverage of recent advances on possibly the next generation of posterior samplers. The final paper, Bayesian experimental design by Rainforth & al., provides a most convincing application of the methods exposed in the earlier papers in that the field of Bayesian design has hugely benefited from the occurrence of such tools to become a prevalent way of designing statistical experiments in real settings.

We feel the future of Bayesian computing is bright! The Monte Carlo revolution of the 1990s continues to be a huge influence on today’s work, and now is complemented by an exciting range of new directions informed by modern machine learning.

Dennis Prangle and Christian P Robert