Archive for Voronoi tesselation

on control variates

Posted in Books, Kids, Statistics, University life with tags , , , , , , , , , , , , on May 27, 2023 by xi'an

A few months ago, I had to write a thesis evaluation of Rémi Leluc’s PhD, which contained several novel Monte Carlo proposals on control variates and importance techniques. For instance, Leluc et al. (Statistics and Computing, 2021) revisits the concept of control variables by adding a perspective of control variable selection using LASSO. This prior selection is relevant since control variables are not necessarily informative about the objective function being integrated and my experience is that the more variables the less reliable the improvement. The remarkable feature of the results is in obtaining explicit and non-asymptotic bounds.

The author obtains a concentration inequality on the error resulting from the use of control variables, under strict assumptions on the variables. The associated numerical experiment illustrates the difficulties of practically implementing these principles due to the number of parameters to calibrate. I found the example of a capture-recapture experiment on ducks (European Dipper) particularly interesting, not only because we had used it in our book but also because it highlights the dependence of estimates on the dominant measure.

Based on a NeurIPS 2022 poster presentation Chapter 3 is devoted to the use of control variables in sequential Monte Carlo, where a sequence of importance functions is constructed based on previous iterations to improve the approximation of the target distribution. Under relatively strong assumptions of importance functions dominating the target distribution (which could generally be achieved by using an increasing fraction of the data in a partial posterior distribution), of sub-Gaussian tails of an intractable distribution’s residual, a concentration inequality is established for the adaptive control variable estimator.

This chapter uses a different family of control variables, based on a Stein operator introduced in Mira et al. (2016). In the case where the target is a mixture in IRd, one of our benchmarks in Cappé et al. (2008), remarkable gains are obtained for relatively high dimensions. While the computational demands of these improvements are not mentioned, the comparison with an MCMC approach (NUTS) based on the same number of particles demonstrates a clear improvement in Bayesian estimation.

Chapter 4 corresponds to a very recent arXival and presents a very original approach to control variate correction by reproducing the interest rate law through an approximation using the closest neighbor (leave-one-out) method. It requires neither control function nor necessarily additional simulation, except for the evaluation of the integral, which is rather remarkable, forming a kind of parallel with the bootstrap. (Any other approximation of the distribution would also be acceptable if available at the same computational cost.) The thesis aims to establish the convergence of the method when integration is performed by a Voronoi tessellation, which leads to an optimal rate of order n-1-2/d for quadratic error (under conditions of integrand regularity). In the alternative where the integral must be evaluated by Monte Carlo, this optimality disappears, unless a massive amount of simulations are used. Numerical illustrations cover SDEs and a Bayesian hierarchical modeling already used in Oates et al. (2017), with massive gain in both cases.

ABC for COVID spread reconstruction

Posted in Books, pictures, Statistics, Travel with tags , , , , , , , , , on December 27, 2021 by xi'an

A recent Nature paper by Jessica Davis et al. (with an assessment by Simon Cauchemez and X from INSERM) reassessed the appearance of COVID in European and American States. Accounting for the massive under-reporting in the early days since there was no testing. The approach is based on a complex dynamic model whose parameters are estimated by an ABC algorithm (the reference being the PLoS article that initiated the ABC Wikipedia page). Results are quite interesting in that the distribution of the entry dates covers a calendar as early as December 2019 in most cases. And a proportion of missed cases as high as 99%.

“As evidence, E, we considered the cumulative number of SARS-CoV-2 cases internationally imported from China up to January 21, 2020″

The model behind remain a classical SLIR model but with a discrete and stochastic dynamical and a geographical compartmentalization based on a Voronoi tessellation centred at airports, commuting intensity and population density. Interventions by local and State authorities are also accounted for. The ABC version is a standard rejection algorithm with distance based on the evidence as quoted above. Which is a form of cdf distance (as in our Wasserstein ABC paper). For the posterior distribution of the IFR,  a second ABC algorithm uses the relative distance between observed and generated deaths (per country). The paper further investigates different introduction sources (countries) before local transmission was established. For instance, China is shown to be the dominant source for the first EU countries impacted by the pandemics such as Italy, UK, Germany, France and Spain. Using a “counterfactual scenario where the surveillance systems of the US states and European countries are imagined to operate at levels able to identify 50% of all imported and locally generated infections”, the authors conclude that

“broadening testing specifications could have considerably slowed the pandemic progression, buying considerable time to prepare mitigation responses.”

Statistics at Bristol [& U]

Posted in pictures, Statistics, Travel, University life with tags , , , , , on September 14, 2021 by xi'an

For the celebration of the recently renovated Fry Building, which I visited in Feb 2019 (my last time in Britain!), the University of Bristol is holding the Fry Conference series, with one dedicated to statistics on 16-17 September 2021. With Peter Green, Arnaud Doucet, and Judith Rousseau among the speakers. It is sadly on-line so does not give one the opportunity to admire the renovated bulding. And the Voronoi sculpture! (And You figures in the title of the conference.)

The Fry Building [Bristol maths]

Posted in Kids, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , on March 7, 2020 by xi'an

While I had heard of Bristol maths moving to the Fry Building for most of the years I visited the department, starting circa 1999, this last trip to Bristol was the opportunity for a first glimpse of the renovated building which has been done beautifully, making it the most amazing maths department I have ever visited.  It is incredibly spacious and luminous (even in one of these rare rainy days when I visited), while certainly contributing to the cohesion and interactions of the whole department. And the choice of the Voronoi structure should not have come as a complete surprise (to me), given Peter Green’s famous contribution to their construction!

lords of the rings

Posted in Books, pictures, Statistics, University life with tags , , , , , , on February 9, 2017 by xi'an

In the 19 Jan 2017 issue of Nature [that I received two weeks later], a paper by Tarnita et al discusses regular vegetation patterns like fairy patterns. While this would seem like an ideal setting for point process modelling, the article does not seem to get into that direction, debating instead between ecological models. Which combines vegetal self-organisation, with subterranean insect competition. Since the paper seems to derive validation of a model by simulation means without producing a single equation, I went and checked the supplementary material attached to this paper. What I gathered from this material is that the system of differential equations used to build this model seems to be extrapolated by seeking parameter values consistent with what is known” rather than estimated as in a statistical model. Given the extreme complexity of the resulting five page model, I am surprised at the low level of validation of the construct, with no visible proof of stationarity of the (stochastic) model thus constructed, and no model assessment in a statistical sense. Of course, a major disclaimer applies: (a) this area does not even border my domains of (relative) expertise and (b) I have not spent much time perusing over the published paper and the attached supplementary material. (Note: This issue of Nature also contains a fascinating review paper by Nielsen et al. on a detailed scenario of human evolutionary history, based on the sequencing of genomes of extinct hominids.)