Archive for rare events

more than mostly MC

Posted in Kids, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , on December 27, 2024 by xi'an

The session of last Friday (and last one of 2024!) proved most interesting, with two (fully) Monte Carlo talks. The first one by Louis Grenioux was about an improvement on diffusion sampler, following a recent arXival by Maxime Noble and co-authors. Which brought me back to the long-standing multimodal challenge in Monte Carlo methods. When all modes of a target distribution are known, even roughly, this is not much of an issue since samplers can be arm-bent into visiting all these modes. But the problem becomes much harder when the location and a fortiori the number of modes are not known. The paper aims at adapting diffusion samplers towards a better exploration of the modes, albeit their location is known. Otherwise, using MCMC as a starting (reference) distribution would risk missing some of them. Which also explains why the authors can rely on the classical Gaussian mixtures proposal as a cheap substitute to neural networks (EBM). Since, within diffusion models, both intermediary distributions and their scores are intractable, they also introduce a variational parametric approximation that can be optimised.

The second talk was given by Guillaume Chennetier, in connection with his recent PhD thesis, developing a form of X-entropy sampling for rare events. As in nuclear plant major accidents. It took me a while to realise that PDMPs were not used as a simulation tool, as in the zigzag sampler and its avatars, The proposal involved creating a graph structure on the space of PDMP trajectories and designing the optimal importance process (yes, the one with zero variance!) using so-called committor functions that modify jump intensity and kernel, in a sequential way reminiscent of X-entropy. The approach recycles past trajectories as Monte Carlo elements if missing an adaptive mixture importance sampling (AMIS!) version that would bring more stability. The talk also included an interesting pointer to the availability of the distribution of the PDMP path, thus treated as a likelihood. (!). The signage at the entrance of the Monte-Carlo (mind the hyphen!) casino also made an appearance, reminding me of our memorable group picture on the same spot, eons ago! (But not of whom took the picture!)

R[are]SS meeting

Posted in Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on September 29, 2024 by xi'an


Yesterday, I happened to be at the right time in the right place, as I was in Warwick for a RSS local section meeting on rare event simulation. (If missing the aurora borealis and the moon eclipse on previous nights!) And hence attended a seminar by Francesca Crucinio in six days!, as she talked about a turnkey approach to unbiased estimation of transforms of a moment, or wlog a mean μ, f(μ). A recent article with Nicolas Chopin (CREST) and Sumeet Singh, where they resort to Taylor expansions to achieve unbiasedness, using the Russian roulette trick to stop the summation from running to infinity. (As it happens, I heard Nicolas talk about this idea in the recent past namely at the ISBA-Fusion Sunday morn at Ca’Foscari.) Using a Taylor expansion is obviously natural and mathematically correct, albeit fraught with potential dangers [imho]:

  • the Taylor expansion involves central moments up to a random order R, which are harder & harder to estimate with increasing orders (i.e., more & more uncertain, with the possibility of infinite variance estimators after a certain order)
  • I did not spot a discussion on the moment estimators, that seems to rely on k iid replicas for the k-th moment
  • a lot of calibration ensues, from the choice of the centre x⁰ to the (artificial) distribution of the stopping value R, to the parameterisation of the random variable attached to the moment μ
  • the paper insists on recycling simulations to stabilise the moment estimators and ensure consistency, as a primary level of Rao-Blackwellisation, but this only applies to the smallest order moments and could be devised in many different ways, with varying computing costs
  • consistency of the estimate is not necessarily needed, as for instance for pseudo-marginal applications
  • as often with Russian roulette, positive quantities may receive negative estimations that are dominated by truncations to the positive real line (and alternating series offer the use of sandwiching estimators)
  • for the above reason, it is not always reasonable to tunnel vision on unbiasedness and alternative estimates like bridge sampling solutions could be integrating towards improving the quality of the estimator (especially since the conditions for finite variance involve unknown quantities)
  • while f-Taylored solutions like harmonic mean estimators for f(x)=1/x are not necessarily a panacea, they could be included in the comparison or as control variates

The first talk by Mathias Rousset was investigating adaptive multilevel sampling, a form of nested sampler, at the theoretical level, while the third talk by Tobias Grafke was a repetition of a talk he gave at the masterclass the interface between computational physics and computational statistics, last April.

rare ABC [webinar impressions]

Posted in Books, Statistics, Travel, University life with tags , , , , , , , on April 28, 2020 by xi'an

A second occurrence of the One World ABC seminar by Ivis Kerama, and Richard Everitt (Warwick U), on their on-going pape with and Tom Thorne, Rare Event ABC-SMC², which is not about rare event simulation but truly about ABC improvement. Building upon a previous paper by Prangle et al. (2018). And also connected with Dennis’ talk a fortnight ago in that it exploits an autoencoder representation of the simulated outcome being H(u,θ). It also reminded me of an earlier talk by Nicolas Chopin.

This approach avoids using summary statistics (but relies on a particular distance) and implements a biased sampling of the u’s to produce outcomes more suited to the observation(s). Almost sounds like a fiducial ABC! Their stopping rule for decreasing the tolerance is to spot an increase in the variance of the likelihood estimates. As the method requires many data generations for a single θ, it only applies in certain settings. The ABC approximation is indeed used as an estimation of likelihood ratio (which makes sense for SMC² but is biased because of ABC). I got slightly confused during Richard’s talk by his using the term of unbiased estimator of the likelihood before I realised he was talking of the ABC posterior. Thanks to both speakers, looking forward the talk by Umberto Picchini in a fortnight (on a joint paper with Richard).

rare events for ABC

Posted in Books, Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , on November 24, 2016 by xi'an

Dennis Prangle, Richard G. Everitt and Theodore Kypraios just arXived a new paper on ABC, aiming at handling high dimensional data with latent variables, thanks to a cascading (or nested) approximation of the probability of a near coincidence between the observed data and the ABC simulated data. The approach amalgamates a rare event simulation method based on SMC, pseudo-marginal Metropolis-Hastings and of course ABC. The rare event is the near coincidence of the observed summary and of a simulated summary. This is so rare that regular ABC is forced to accept not so near coincidences. Especially as the dimension increases.  I mentioned nested above purposedly because I find that the rare event simulation method of Cérou et al. (2012) has a nested sampling flavour, in that each move of the particle system (in the sample space) is done according to a constrained MCMC move. Constraint derived from the distance between observed and simulated samples. Finding an efficient move of that kind may prove difficult or impossible. The authors opt for a slice sampler, proposed by Murray and Graham (2016), however they assume that the distribution of the latent variables is uniform over a unit hypercube, an assumption I do not fully understand. For the pseudo-marginal aspect, note that while the approach produces a better and faster evaluation of the likelihood, it remains an ABC likelihood and not the original likelihood. Because the estimate of the ABC likelihood is monotonic in the number of terms, a proposal can be terminated earlier without inducing a bias in the method.

Lake Louise, Banff National Park, March 21, 2012This is certainly an innovative approach of clear interest and I hope we will discuss it at length at our BIRS ABC 15w5025 workshop next February. At this stage of light reading, I am slightly overwhelmed by the combination of so many computational techniques altogether towards a single algorithm. The authors argue there is very little calibration involved, but so many steps have to depend on as many configuration choices.

an extension of nested sampling

Posted in Books, Statistics, University life with tags , , , , , , , on December 16, 2014 by xi'an

I was reading [in the Paris métro] Hastings-Metropolis algorithm on Markov chains for small-probability estimation, arXived a few weeks ago by François Bachoc, Lionel Lenôtre, and Achref Bachouch, when I came upon their first algorithm that reminded me much of nested sampling: the following was proposed by Guyader et al. in 2011,

To approximate a tail probability P(H(X)>h),

  • start from an iid sample of size N from the reference distribution;
  • at each iteration m, select the point x with the smallest H(x)=ξ and replace it with a new point y simulated under the constraint H(y)≥ξ;
  • stop when all points in the sample are such that H(X)>h;
  • take

\left(1-\dfrac{1}{N}\right)^{m-1}

as the unbiased estimator of P(H(X)>h).

Hence, except for the stopping rule, this is the same implementation as nested sampling. Furthermore, Guyader et al. (2011) also take advantage of the bested sampling fact that, if direct simulation under the constraint H(y)≥ξ is infeasible, simulating via one single step of a Metropolis-Hastings algorithm is as valid as direct simulation. (I could not access the paper, but the reference list of Guyader et al. (2011) includes both original papers by John Skilling, so the connection must be made in the paper.) What I find most interesting in this algorithm is that it even achieves unbiasedness (even in the MCMC case!).