Archive for Comptes Rendus de l’Académie des Sciences

inverse probability weighting

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , on May 4, 2026 by xi'an

Quite recently, Jyotishka Datta and Nick Polson published a fairly interesting [imho] paper in The New England Journal of Statistics in Data Science, entitled Inverse Probability Weighting: From Survey Sampling to Evidence Estimation that (obviously) caters to my own interests! They bring three threads together. First, they recall the long debate between using [normalized] Horvitz–Thompson and [self-normalized] Hájek estimators in survey sampling, pointing out that the latter is “usually the better estimator, despite estimation of an a priori known quantity” (citing from Särndal & al., 2003). Which is also my experience with importance sampling, as in this 1995 Note aux Comptes Rendus with George. The mathematical paradox of “estimating” a constant is central to other advances in the area, like noise-contrastive estimation à la Gutmann & Hyvärinen (2005) or the measure estimation of Kong & al. (2003). Even more interestingly, Datta & Polson consider there is a link with the inconsistent Bayesian (counter)example of Larry Wasserman and Jamie Robins, where the censoring probability increases with the value of the parameter of interest, paradox in which Chris Sims also got involved. (I was unaware that he had passed away last month.) And, lo and behold!, with the Stein “paradox” of my PhD years (and beyond).

Their central argument stands with the missing data link between Horvitz-Thompson survey sampling and Monte Carlo integration. (With a reference to our Riemann sum papers with Anne Philippe!, making me realise the authors had recently published an extension on that idea.) This reminds me very much of the missing measure approach of Kong et al. (2003). (Actually the reference appears in the final discussion.) The authors go over several paradoxes like Basu’s circus estimate (1988), Larry’s inconsistent Bayes estimate (2004), where Horvitz–Thompson performs nicely under compactness assumptions, the Bayesian answers (which include nested sampling even though I do not see the connection). Especially Li’s (2010) solution.  The attached numerical experiment displays a consistent underperformance of the Horvitz-Thompson estimator, in contrast with the theory… 

In conclusion, while enjoying very much revisiting so many examples and papers I came across in the past decades, I remain somewhat puzzled by the lack of overall message.

 

gentle importance sampling

Posted in Books, pictures, Statistics with tags , , , , , , , , , , , , on February 24, 2025 by xi'an

A new (and gentle!) survey by Luca Martino! And by Fernando Llorente. On importance sampling, with coverage of normalised and self-normalised versions. And their usage in different configurations (one vs several integrals, one vs several families of distributions). Some points relating to earlier remarks or musing of mine’s:

  • the fact that the optimal importance function does not lead to a zero variance importance estimator when the integrand f is not of constant sign (p.7) can be cancelled by first decomposing f as f⁺-f⁻, since both allow for a zero variance importance estimator, if formally requiring two different samples (of size zero!), a trick considered later on p.18 and repeated for the ratio in self-normalised importance (p.19)
  • the special case when the integrand f is constant is not of practical interest but relevant for checking properties of different estimators. For instance, this case allowed George and myself to spot a mistake in an early importance paper. In the same volume of the Comptes Rendus as an early paper of Lions and Villani.
  • the remark that self-normalised (SNIS) importance sampling can prove more efficient than (properly normalised) importance sampling, although the property that SNIS is always bounded should not be seen as a major point given that it is simply due to using a finite sample and hence a finite set of images of f
  • the case of integrals involving several target pdfs or several integrands is not necessarily of major interest if simulating different samples for each unidimensional integral can be implemented (again formally leading to zero variance at no cost)
  • the issue of merging several estimators in an optimal way is briefly mentioned in §5.4, a challenge Victor Elvira and I have been approaching over the past years, if not yet concluding satisfactorily (mea culpa)
  • when replacing the target with a noisy estimate (p.22), the fact that this estimate must be normalised is correct, but pales against the impact of using this estimate, which may prove catastrophic. And unbiasedness is not particularly crucially important in this setup for the same reason
  • the section on evidence approximation (§7) is more standard, with the harmonic mean estimator being called reverse importance sampling, which brings us to the “elephant in the room”, namely that
  • the issue of infinite variance of some importance sampling estimators is not directly covered (except once in §8, p.34), thus perceiving importance sampling as a variance reduction method being somewhat misleading (unless the authors consider solely the optimal importance function, which is rarely of practical use)

The paper concludes with an interesting notion that

“we suggest the analysis of the relevant connection between importance sampling and contrastive learning Gutmann and Hyvärinen (2012)”

that I also have been pointing out for a while. All in all, a useful summing-up that I will likely suggest to my students.

my neighbour who invented the gradient descent from the top of his hill

Posted in Books, pictures, University life with tags , , , , , , on September 6, 2024 by xi'an

“It will suffice therefore, either to resolved this last equation, or at least to attribute to θ a sufficiently small value, in order to obtain a new value of u inferior to u. If the new value of u is not a minimum, one will be able to deduce, by operating always in the same manner, a third still smaller value; and, by continuing thus, one will find successively some values of u more and more small, which will converge toward a minimum value of u. If the function u, which is supposed not at all to admit negative values, offers some null values, there will be able always to be determined by the preceding method, provided that one chooses conveniently the values of x, y, z, …” Augustin Cauchy, CRAS, 25, 536–538, 1847 [translation of Richard J. Pulskamp, 2010]

Surprisingly, I was not aware till a few days ago that Augustin Cauchy had made the first proposal of the gradient method to solve a minimisation problem. Which he published in a fairly vague(witness the above quote) paper in the Comptes Rendus de l’Académie des Sciences. Cauchy lived most of his (adult) life in my town of Sceaux, with his house still standing near the town centre, and now part of the nearby Marie Curie high school (named after another famous resident!).

Pitman medal for Kerrie Mengersen

Posted in pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , on December 20, 2016 by xi'an

6831250-3x2-700x467My friend and co-author of many years, Kerrie Mengersen, just received the 2016 Pitman Medal, which is the prize of the Statistical Society of Australia. Congratulations to Kerrie for a well-deserved reward of her massive contributions to Australian, Bayesian, computational, modelling statistics, and to data science as a whole. (In case you wonder about the picture above, she has not yet lost the medal, but is instead looking for jaguars in the Amazon.)

This medal is named after EJG Pitman, Australian probabilist and statistician, whose name is attached to an estimator, a lemma, a measure of efficiency, a test, and a measure of comparison between estimators. His estimator is the best equivariant (or invariant) estimator, which can be expressed as a Bayes estimator under the relevant right Haar measure, despite having no Bayesian motivation to start with. His lemma is the Pitman-Koopman-Darmois lemma, which states that outside exponential families, sufficient is essentially useless (except for exotic distributions like the Uniform distributions). Darmois published the result first in 1935, but in French in the Comptes Rendus de l’Académie des Sciences. And the measure of comparison is Pitman nearness or closeness, on which I wrote a paper with my friends Gene Hwang and Bill Strawderman, paper that we thought was the final paper on the measure as it was pointing out several majors deficiencies with this concept. But the literature continued to grow after that..!

une vie brève [book review]

Posted in Books, Kids, pictures, University life with tags , , , , , , , on April 30, 2016 by xi'an

This short book is about the equally short life (une vie brève) of the young mathematician Maurice Audin, killed or executed by French special forces (Massu’s paratroopers) in Algiers during the Algerian liberation war. Maurice Audin was 25 when he died and the circumstances of his death remain unknown, since the French army never acknowledged this death and never returned his body to his family, but he presumably died under torture. He was a member of the Algerian communist party which had then been outlawed by the French authorities for supporting Algerian independence. Maurice Audin was arrested on June 11, 1957 for hiding a fugitive and he died in the following days… The book is written by his daughter, Michèle Audin, also a mathematician, and a writer of several novels around mathematics and mathematicians. It does not dwell on the death since so little is known but rather reconstructs the life of Maurice Audin from bits and pieces, family memories, school archives, a few pictures, some grocery bills of the Audin family… The style of Michèle Audin is quite peculiar, almost like written thoughts or half-thoughts at times, with a sort of surgical distanciation that makes the book both strong and touching. Maurice Audin wrote several papers in les Comptes Rendus de l’Académie des Sciences [the French PNAS] but did not live long enough to defend his thesis, which was presented by Laurent Schwartz the following year and defended in absentia… The French State never acknowledged its responsability in Audin’s death. (Another book on this death is L’Affaire Audin by the historian Pierre-Vidal Naquet, which appeared in 1958.)