Archive for CRAS

gentle importance sampling

Posted in Books, pictures, Statistics with tags , , , , , , , , , , , , on February 24, 2025 by xi'an

A new (and gentle!) survey by Luca Martino! And by Fernando Llorente. On importance sampling, with coverage of normalised and self-normalised versions. And their usage in different configurations (one vs several integrals, one vs several families of distributions). Some points relating to earlier remarks or musing of mine’s:

  • the fact that the optimal importance function does not lead to a zero variance importance estimator when the integrand f is not of constant sign (p.7) can be cancelled by first decomposing f as f⁺-f⁻, since both allow for a zero variance importance estimator, if formally requiring two different samples (of size zero!), a trick considered later on p.18 and repeated for the ratio in self-normalised importance (p.19)
  • the special case when the integrand f is constant is not of practical interest but relevant for checking properties of different estimators. For instance, this case allowed George and myself to spot a mistake in an early importance paper. In the same volume of the Comptes Rendus as an early paper of Lions and Villani.
  • the remark that self-normalised (SNIS) importance sampling can prove more efficient than (properly normalised) importance sampling, although the property that SNIS is always bounded should not be seen as a major point given that it is simply due to using a finite sample and hence a finite set of images of f
  • the case of integrals involving several target pdfs or several integrands is not necessarily of major interest if simulating different samples for each unidimensional integral can be implemented (again formally leading to zero variance at no cost)
  • the issue of merging several estimators in an optimal way is briefly mentioned in §5.4, a challenge Victor Elvira and I have been approaching over the past years, if not yet concluding satisfactorily (mea culpa)
  • when replacing the target with a noisy estimate (p.22), the fact that this estimate must be normalised is correct, but pales against the impact of using this estimate, which may prove catastrophic. And unbiasedness is not particularly crucially important in this setup for the same reason
  • the section on evidence approximation (§7) is more standard, with the harmonic mean estimator being called reverse importance sampling, which brings us to the “elephant in the room”, namely that
  • the issue of infinite variance of some importance sampling estimators is not directly covered (except once in §8, p.34), thus perceiving importance sampling as a variance reduction method being somewhat misleading (unless the authors consider solely the optimal importance function, which is rarely of practical use)

The paper concludes with an interesting notion that

“we suggest the analysis of the relevant connection between importance sampling and contrastive learning Gutmann and Hyvärinen (2012)”

that I also have been pointing out for a while. All in all, a useful summing-up that I will likely suggest to my students.

my neighbour who invented the gradient descent from the top of his hill

Posted in Books, pictures, University life with tags , , , , , , on September 6, 2024 by xi'an

“It will suffice therefore, either to resolved this last equation, or at least to attribute to θ a sufficiently small value, in order to obtain a new value of u inferior to u. If the new value of u is not a minimum, one will be able to deduce, by operating always in the same manner, a third still smaller value; and, by continuing thus, one will find successively some values of u more and more small, which will converge toward a minimum value of u. If the function u, which is supposed not at all to admit negative values, offers some null values, there will be able always to be determined by the preceding method, provided that one chooses conveniently the values of x, y, z, …” Augustin Cauchy, CRAS, 25, 536–538, 1847 [translation of Richard J. Pulskamp, 2010]

Surprisingly, I was not aware till a few days ago that Augustin Cauchy had made the first proposal of the gradient method to solve a minimisation problem. Which he published in a fairly vague(witness the above quote) paper in the Comptes Rendus de l’Académie des Sciences. Cauchy lived most of his (adult) life in my town of Sceaux, with his house still standing near the town centre, and now part of the nearby Marie Curie high school (named after another famous resident!).