Archive for hit and run algorithm

Choice [book review]

Posted in Books, Kids, pictures, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , on August 31, 2025 by xi'an

I first got attracted by this book thanks to its beautiful cover (in a Seattle bookstore last year)! The book is an aggregate of three stories, loosely related around the themes of altruism and disastrous good intentions. I did not like the first story, about the disintegration of a gay couple by ways of (OCD) psychiatric issues as well as an increasing radicalism towards Ayush’s societal choices, with some shocking episodes as when he shows illegal filming from pig abattoirs to their young children. I saw some worth in the second story, when a classic academic gets obsessed with the migration of a Sudanese child-soldier to the point of obsession, removing herself from her job and civic duties, as in not reporting a possible hit & run, and eventually gifting a kidney to the refugee’s brother. As a dubious reparation for her grand-parents’ involvement in the British Raj colonialism. I mostly enjoyed the third one, where Sabita, a rural Bengali or Bangladeshi woman, receives a cow from experimental economists (in the spirit of Esther Duflo’s school), a stupendous gift that is slowly unraveling her family and her precarious finances. However, my overall impression is one of an overly ideological posture, tending to caricatures (esp. of academics) and compartmentalisation of individuals, to the detriment of the book per se and to the depth of its characters. The last story is further strongly condescending towards the main character who proves unable to manage the cow (and her family), again forcing the trait for the sake of the political argument. The Guardian is more appreciative of the book, however.

All about that [Bayes] seminar [24 Jan]

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on January 13, 2025 by xi'an

The next All about that (Bayes) seminar will take place on Friday 24 Jan at SCAI, on the Jussieu campus, with the following talks. (Appearances to the contrary, I was not in the least involved in the program!)

13h30 – 14h30 Joshua Bon (OCEAN, Université Paris Dauphine) – Bayesian score calibration for approximate models

 Scientists continue to develop increasingly complex mechanistic models to reflect their knowledge more realistically. Statistical inference using these models can be challenging since the corresponding likelihood function is often intractable and model simulation may be computationally burdensome. Fortunately, in many of these situations, it is possible to adopt a surrogate model or approximate likelihood function. It may be convenient to conduct Bayesian inference directly with the surrogate, but this can result in bias and poor uncertainty quantification. In this paper (https://arxiv.org/abs/2211.05357) we propose a new method for adjusting approximate posterior samples to reduce bias and produce more accurate uncertainty quantification. We do this by optimizing a transform of the approximate posterior that maximizes a scoring rule. Our approach requires only a (fixed) small number of complex model simulations and is numerically stable. We demonstrate beneficial corrections to several approximate posteriors using our method on several examples of increasing complexity.

14h30 – 15h30 Giacomo Zanella (Bocconi University) – Entropy contraction of the Gibbs sampler under log-concavity

In this talk I will present recent work (https://arxiv.org/abs/2410.00858) on the non-asymptotic analysis of the Gibbs sampler, a classical and popular MCMC algorithm for sampling. In particular, under the assumption that the probability measure π of interest is strongly log-concave, we show that the random scan Gibbs sampler contracts in relative entropy, and provide a sharp characterization of the associated contraction rate. The result implies that, under appropriate conditions, the number of full evaluations of π required for the Gibbs sampler to converge is independent of the dimension. If time permits, I will also discuss connections and applications of the above results to the problem of zero-order parallel sampling, as well as extensions to Hit-and-Run and Metropolis-within-Gibbs.

Based on joint work with Filippo Ascolani and Hugo Lavenant.

16h00 – 17h00 Paul Bastide (Université Paris Cité) – Goodness of Fit for Bayesian Generative Models with Applications in Population Genetics

In population genetics, inference about intractable likelihood models is common, and simulation methods, including Approximate Bayesian Computation (ABC) and Simulation-Based Inference (SBI), are essential. ABC/SBI methods work by simulating instrumental data sets of the models under study and comparing them with the observed data set y⁰. Advanced machine learning tools are used for tasks such as model selection and parameter inference. The present work focuses on model criticism. This type of analysis, called goodness of fit (GoF), is important for model validation. It can also be used for model pruning when the number of candidates to be considered is excessive, especially in the context where data simulation is expensive. We introduce two new GoF tests based on the local outlier factor (LOF), an indicator that was initially defined for outlier and novelty detection. We test whether y⁰ is distributed from the prior predictive distribution (pre-inference GoF) and whether there is a parameter value such that y⁰ is distributed from the likelihood with that value (post-inference GoF).  We evaluate the performance of our two GoF tests on simulated datasets from three different model settings of varying complexity, and on a dataset of single nucleotide polymorphism (SNP) markers for the evaluation of complex evolutionary scenarios of modern human populations.

Joint work with Guillaume Le Mailloux, Jean-Michel Marin and Arnaud Estoup.

skipping sampler

Posted in Books, Statistics, University life with tags , , , , on June 13, 2019 by xi'an

“The Skipping Sampler is an adaptation of the MH algorithm designed to sample from targets which have areas of zero density. It ‘skips’ across such areas, much as a flat stone can skip or skim repeatedly across the surface of water.”

An interesting challenge is simulating from a density restricted to a set C when little is known about C, apart from a mean to check whether or not a given value x is in C or not. John Moriarty, Jure Vogrinc (University of Warwick), and Alessandro Zocca make a new proposal to address this problem in a recently arXived paper. Which somewhat reminded me of the delayed rejection methods proposed by Antonietta Mira. And of our pinball sampler.

The paper spends a large amount of space about transferring from the Euclidean representation of the symmetric proposal density q to its polar representation. Which is rather trivial, but brings the questions of the efficient polar proposals  and of selecting the right type of Euclidean distance for the intended target. The method proposed therein is to select a direction first and keep skipping step by step in that direction until the set C is met again (re-entered). Or until a stopping (halting) boundary has been hit. This makes for a more complex proposal than usual but somewhat surprisingly the symmetry in q is sufficient to make the acceptance probability only depend on the target density.

While the convergence is properly established, I wonder at the practicality of the approach when compared with a regular random walk Metropolis algorithm in that both require a scaling to the jump that relates to the support of the target. Neither too small nor too large. If the set C is that unknown that only local (in or out) information is available, scaling of the jumps (and of the stopping rule) may prove problematic. In equivalent ways for both samplers. In a completely blind exploration, sequential (or population) Monte Carlo would seem more appropriate, at least to learn about the scale of jumps and location of the set C. If this set is defined as an intersection of constraints, a tempered (and sequential) solution would be helpful.  When checking the appurtenance to C becomes a computational challenge, more advance schemes have to be constructed, I would think.