Archive for Gaussian mixture

OWABI@BioInference2025 [29 May]

Posted in Mountains, Statistics, Travel, University life with tags , , , , , , , , , , , , , , on May 13, 2025 by xi'an

The next OWABI webinar is going to be quite special, consisting of two selected talks livestreamed from BioInference 2025, a conference on mathematical modelling and inference on (broadly speaking) biological system, taking place in Bardonecchia, Piedmont. The talks will take place on 29 May, 11am CEST (10am BST). The talks will be streamed on the OWABI MS Team Channel as usual.

1st OWABI Talk: 10-10.30am UK time

Speaker: Andrew Golightly (Durham University)

Title: Accelerating Bayesian inference for stochastic epidemic models using incidence data

Abstract: This work considers the case of performing Bayesian inference for stochastic epidemic compartment models, using incomplete time course data consisting of incidence counts that are either the number of new infections or removals in time intervals of fixed length. The most natural Markov jump process representation of the model is eschewed for reasons of computational efficiency, and replaced by a stochastic differential equation representation. This is further approximated to give a tractable Gaussian process, that is, the linear noise approximation (LNA). Unless the observation model linking the LNA to data is both linear and Gaussian, the observed data likelihood remains intractable. Unlike previous approaches that use the LNA in this setting, two approaches for marginalising over the latent process are considered: a correlated pseudo-marginal method and analytic marginalisation via a Gaussian approximation of the noise model. These approaches are compared using synthetic data with the best performing method applied to real data consisting of removal incidence of Oak Processionary moth nests in Richmond Park, London.

2nd OWABI Talk: 10.30-11am

Speaker: Henrik Häggström (Chalmers University)

Title: Simulation-based inference for stochastic nonlinear mixed-effects models with applications in systems biology

Abstract: We propose a novel methodology for Bayesian inference in hierarchical mixed-effects models. By building on our work, we construct a simulation-based inference (SBI) framework that is highly scalable, where amortized approximations to the likelihood and the parameters posterior are first obtained, and these are rapidly refined for each individual dataset, to ultimately approximate the parameters posterior across many individuals. Unlike the current state-of-art SBI methods, which use neural networks, our approximations are expressed via Gaussian mixture models, leading to easily trainable, parsimonious yet expressive surrogate models of both the likelihood function and the posterior distribution. The methodology is exemplified via stochastic differential equation mixed-effects models to describe translation kinetics after mRNA transfection, however the methodology is general and can accommodate other types of stochastic and deterministic models. We compare our approximate inference with exact pseudomarginal inference and show that our methodology is fast and competitive.

more than mostly MC

Posted in Kids, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , on December 27, 2024 by xi'an

The session of last Friday (and last one of 2024!) proved most interesting, with two (fully) Monte Carlo talks. The first one by Louis Grenioux was about an improvement on diffusion sampler, following a recent arXival by Maxime Noble and co-authors. Which brought me back to the long-standing multimodal challenge in Monte Carlo methods. When all modes of a target distribution are known, even roughly, this is not much of an issue since samplers can be arm-bent into visiting all these modes. But the problem becomes much harder when the location and a fortiori the number of modes are not known. The paper aims at adapting diffusion samplers towards a better exploration of the modes, albeit their location is known. Otherwise, using MCMC as a starting (reference) distribution would risk missing some of them. Which also explains why the authors can rely on the classical Gaussian mixtures proposal as a cheap substitute to neural networks (EBM). Since, within diffusion models, both intermediary distributions and their scores are intractable, they also introduce a variational parametric approximation that can be optimised.

The second talk was given by Guillaume Chennetier, in connection with his recent PhD thesis, developing a form of X-entropy sampling for rare events. As in nuclear plant major accidents. It took me a while to realise that PDMPs were not used as a simulation tool, as in the zigzag sampler and its avatars, The proposal involved creating a graph structure on the space of PDMP trajectories and designing the optimal importance process (yes, the one with zero variance!) using so-called committor functions that modify jump intensity and kernel, in a sequential way reminiscent of X-entropy. The approach recycles past trajectories as Monte Carlo elements if missing an adaptive mixture importance sampling (AMIS!) version that would bring more stability. The talk also included an interesting pointer to the availability of the distribution of the PDMP path, thus treated as a likelihood. (!). The signage at the entrance of the Monte-Carlo (mind the hyphen!) casino also made an appearance, reminding me of our memorable group picture on the same spot, eons ago! (But not of whom took the picture!)

manifold learning [BNP Seminar, 11/01/23]

Posted in Books, Statistics, University life with tags , , , , , , , , on January 9, 2023 by xi'an

An incoming BNP webinar on Zoom by Judith Rousseau and Paul Rosa (U of Oxford), on 11 January at 1700 Greenwich time:

Bayesian nonparametric manifold learning

In high dimensions it is common to assume that the data have a lower dimensional structure. We consider two types of low dimensional structure: in the first part the data is assumed to be concentrated near an unknown low dimensional manifold, in the second case it is assumed to be possibly concentrated on an unknown manifold. In both cases neither the manifold nor the density is known. Atypical example is for noisy observations on an unknown low dimensional manifold.

We first consider a family of Bayesian nonparametric density estimators based on location – scale Gaussian mixture priors and we study the asymptotic properties of the posterior distribution. Our work shows in particular that non conjuguate location-scale Gaussian mixture models can adapt to complex geometries and spatially varying regularity when the density is supported near a low dimensional manifold.

In the second part of the talk we will consider also the case where the distribution is supported on a low dimensional manifold. In this non dominated model,we study different types of posterior contraction rates: Wasserstein and L_1(\mu_\mathcal{M}) where \mu_\mathcal{M} is the Haussdorff measure on the manifold \mathcal{M} supporting the density. Some more generic results on Wasserstein contraction rates are also discussed.

 

scale matters [maths as well]

Posted in pictures, R, Statistics with tags , , , , , , , , on June 2, 2021 by xi'an

A question from X validated on why an independent Metropolis sampler of a three component Normal mixture based on a single Normal proposal was failing to recover the said mixture…

When looking at the OP’s R code, I did not notice anything amiss at first glance (I was about to drive back from Annecy, hence did not look too closely) and reran the attached code with a larger variance in the proposal, which returned the above picture for the MCMC sample, close enough (?) to the target. Later, from home, I checked the code further and noticed that the Metropolis ratio was only using the ratio of the targets. Dividing by the ratio of the proposals made a significant (?) to the representation of the target.

More interestingly, the OP was fundamentally confused between independent and random-walk Rosenbluth algorithms, from using the wrong ratio to aiming at the wrong scale factor and average acceptance ratio, and furthermore challenged by the very notion of Hessian matrix, which is often suggested as a default scale.

parallel tempering on optimised paths

Posted in Statistics with tags , , , , , , , , , , , , , , , on May 20, 2021 by xi'an


Saifuddin Syed, Vittorio Romaniello, Trevor Campbell, and Alexandre Bouchard-Côté, whom I met and discussed with on my “last” trip to UBC, on December 2019, just arXived a paper on parallel tempering (PT), making the choice of tempering path an optimisation problem. They address the touchy issue of designing a sequence of tempered targets when the starting distribution π⁰, eg the prior, and the final distribution π¹, eg the posterior, are hugely different, eg almost singular.

“…theoretical analysis of reversible variants of PT has shown that adding too many intermediate chains can actually deteriorate performance (…) [while] on non reversible regime adding more chains is guaranteed to improve performances.”

The above applies to geometric combinations of π⁰ and π¹. Which “suffers from an arbitrarily suboptimal global communication barrier“, according to the authors (although the counterexample is not completely convincing since π⁰ and π¹ share the same variance). They propose a more non-linear form of tempering with constraints on the dependence of the powers on the temperature t∈(0,1).  Defining the global communication barrier as an average over temperatures of the rejection rate, the path characteristics (e.g., the coefficients of a spline function) can then be optimised in terms of this objective. And the temperature schedule is derived from the fact that the non-asymptotic round trip rate is maximized when the rejection rates are all equal. (As a side item, the technique exposed in the earlier tempering paper by Syed et al. was recently exploited for a night high resolution imaging of a black hole from the M87 galaxy.)