Archive for variance reduction

Monte Carlo with infinite variances [a surveyal guide]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , on January 14, 2026 by xi'an

Watch out!, Reiichiro Kawai has just published a survey on infinite variance Monte Carlo methods in Probability Surveys, which is most welcomed as this issue is customarily ignored by both the literature and the practitioners. Radford Neal‘s warning about the dangers of using the harmonic mean estimator of the evidence (as in Newton and Raftery 1996) is an illustration that remains pertinent to this day. In that sense, the survey relates to specific, earlier if recent attempts, such as Chatterjee and Diaconis (2015) or Vehtari et al (2015), with its Pareto correction.

In its recapitulation of the basics of Monte Carlo (closely corresponding to my own introduction of the topic in undergraduate classes), the paper indicates that the consistency of the variance estimator is enough to replace the true variance with its estimator and maintain the CLT. I have often if vaguely wondered at the impact (if any) a variance estimator with (itself) an infinite variance would have. A note to this effect appears at the end of Section 1.2. While being involved from the start, importance sampling has to wait till section 3.2 to be formally introduced. It is also interesting to note that the original result on the optimal importance variance being zero when the integrand is always positive (or negative) is extended here, by noting that a zero variance estimator can always be found by breaking the integrand f into its positive and negative parts, and using now two single samples for the respective integrals. I thus find Example 6 rather unhelpful, even though the entire literature contains such examples with no added value of formal optimal importance samplers. A comment at the end of Example 6 is opens the door to a short discussion of reparametrisation in simulation, a topic rarely discussed in the literature. The use of Rao-Blackwellization as a variance reduction technique that is open to switching from infinite to finite variance, is emphasised as well in Section 2.1.

In relation with a recent musing of mine during a seminar in Warwick, the novel part in the survey on the limited usefulness of control variate is of interest, even though one could predict that linear regression is not doing very well in infinite variance environments. Examples 8 and 9 are most helpful in this respect. It is similarly revealing if unsurprising that basic antithetic variables do not help. The warning about detecting or failing to detect infinite variance situations is well-received.

While theoretically correct, the final section about truncation limit is more exploratory, in that truncation can produce biased answers, whose magnitude is not assessed within the experiment.

repulsive sampling

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , on January 31, 2024 by xi'an

After a long absence from the monthly Séminaire Parisien de Statistique I attended one today at IHP, including a talk by Diala Hawat on repelled point processes for numerical integration by Hawat et al. The goal is to get (and prove) a universal variance improvement for numerical integration by applying a form of determinantal processes to initial simulations, as eg iid (Poisson process) sampling (without accounting for the O(N²) cost in moving these points). The repelled points are obtain by a single (why single?) move based on a force function (as shown in the slide below), inspired by a Coulomb potential (in the sense that said move appears as one gradient step along the potential). Which reminded me of the pinball sampler, even though the inverse norm was just there to create infinite repulsion near each point. A surprising feature of this repelling step is that it even modifies a (QMC) Sobol process with also an (empirical) improvement in the variance. I wonder if one could construct an MCMC algorithm that would target a joint distribution, maybe via a copula representation, maybe via an equivalent version of HMC.


As an aside, the Bakhvalov results on the existence of a worst case integrand for any deterministic or random sequence (see top slide) made me wonder what the shape of this worst case function is, esp. for a QMC sequence (eg, Sobol). And whether or not they are of any relevance as a counterfactor to the optimal importance functions.

simulating Gumbel’s bivariate exponential distribution

Posted in Books, Kids, R, Statistics with tags , , , , , , , , , , on January 14, 2024 by xi'an

A challenge interesting enough for a sunny New Year morn, found on X validated, namely the simulation of a bivariate exponential distribution proposed by Gumbel in 1960, with density over the positive quadrant in IR²

{}^{ [(\lambda_2+rx_1)(\lambda_1+rx_2)-r]\exp[-(\lambda_1x_1+\lambda_2x_2+rx_1x_2)]}

Although there exists a direct approach based on the fact that the marginals are Exponential distributions and the conditionals signed mixtures of Gamma distributions, an accept-reject algorithm is also available for the pair, with a dominating density representing a genuine mixture of four Gammas, when omitting the X product in the exponential and the negative r in the first term. The efficiency of this accept-reject algorithm is high for r small. However, and in more direct connection with the original question, using this approach to integrate the function equal to the product of the pair, as considered in the original paper of Gumbel, is much less efficient than seeking a quasi-optimal importance function, since this importance function is yet another mixture of four Gammas that produces a much reduced variance at a cheaper cost!

revisiting the balance heuristic

Posted in Statistics with tags , , , , , , , on October 24, 2019 by xi'an

Last August, Felipe Medina-Aguayo (a former student at Warwick) and Richard Everitt (who has now joined Warwick) arXived a paper on multiple importance sampling (for normalising constants) that goes “exploring some improvements and variations of the balance heuristic via a novel extended-space representation of the estimator, leading to straightforward annealing schemes for variance reduction purposes”, with the interesting side remark that Rao-Blackwellisation may prove sub-optimal when there are many terms in the proposal family, in the sense that not every term in the mixture gets sampled. As already noticed by Victor Elvira and co-authors, getting rid of the components that are not used being an improvement without inducing a bias. The paper also notices that the loss due to using sample sizes rather than expected sample sizes is of second order, compared with the variance of the compared estimators. It further relates to a completion or auxiliary perspective that reminds me of the approaches we adopted in the population Monte Carlo papers and in the vanilla Rao-Blackwellisation paper. But it somewhat diverges from this literature when entering a simulated annealing perspective, in that the importance distributions it considers are freely chosen as powers of a generic target. It is quite surprising that, despite the normalising weights being unknown, a simulated annealing approach produces an unbiased estimator of the initial normalising constant. While another surprise therein is that the extended target associated to their balance heuristic does not admit the right density as marginal but preserves the same normalising constant… (This paper will be presented at BayesComp 2020.)

ABC by QMC

Posted in Books, Kids, Statistics, University life with tags , , , , , , , , , , on November 5, 2018 by xi'an

A paper by Alexander Buchholz (CREST) and Nicolas Chopin (CREST) on quasi-Monte Carlo methods for ABC is going to appear in the Journal of Computational and Graphical Statistics. I had missed the opportunity when it was posted on arXiv and only became aware of the paper’s contents when I reviewed Alexander’s thesis for the doctoral school. The fact that the parameters are simulated (in ABC) from a prior that is quite generally a standard distribution while the pseudo-observations are simulated from a complex distribution (associated with the intractability of the likelihood function) means that the use of quasi-Monte Carlo sequences is in general only possible for the first part.

The ABC context studied there is close to the original version of ABC rejection scheme [as opposed to SMC and importance versions], the main difference standing with the use of M pseudo-observations instead of one (of the same size as the initial data). This repeated version has been discussed and abandoned in a strict Monte Carlo framework in favor of M=1 as it increases the overall variance, but the paper uses this version to show that the multiplication of pseudo-observations in a quasi-Monte Carlo framework does not increase the variance of the estimator. (Since the variance apparently remains constant when taking into account the generation time of the pseudo-data, we can however dispute the interest of this multiplication, except to produce a constant variance estimator, for some targets, or to be used for convergence assessment.) L The article also covers the bias correction solution of Lee and Latuszyǹski (2014).

Due to the simultaneous presence of pseudo-random and quasi-random sequences in the approximations, the authors use the notion of mixed sequences, for which they extend a one-dimension central limit theorem. The paper focus on the estimation of Z(ε), the normalization constant of the ABC density, ie the predictive probability of accepting a simulation which can be estimated at a speed of O(N⁻¹) where N is the number of QMC simulations, is a wee bit puzzling as I cannot figure the relevance of this constant (function of ε), especially since the result does not seem to generalize directly to other ABC estimators.

A second half of the paper considers a sequential version of ABC, as in ABC-SMC and ABC-PMC, where the proposal distribution is there  based on a Normal mixture with a small number of components, estimated from the (particle) sample of the previous iteration. Even though efficient techniques for estimating this mixture are available, this innovative step requires a calculation time that should be taken into account in the comparisons. The construction of a decreasing sequence of tolerances ε seems also pushed beyond and below what a sequential approach like that of Del Moral, Doucet and Jasra (2012) would produce, it seems with the justification to always prefer the lower tolerances. This is not necessarily the case, as recent articles by Li and Fearnhead (2018a, 2018b) and ours have shown (Frazier et al., 2018). Overall, since ABC methods are large consumers of simulation, it is interesting to see how the contribution of QMC sequences results in the reduction of variance and to hope to see appropriate packages added for standard distributions. However, since the most consuming part of the algorithm is due to the simulation of the pseudo-data, in most cases, it would seem that the most relevant focus should be on QMC add-ons on this part, which may be feasible for models with a huge number of standard auxiliary variables as for instance in population evolution.