Archive for Jaroslav Hájek

inverse probability weighting

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , on May 4, 2026 by xi'an

Quite recently, Jyotishka Datta and Nick Polson published a fairly interesting [imho] paper in The New England Journal of Statistics in Data Science, entitled Inverse Probability Weighting: From Survey Sampling to Evidence Estimation that (obviously) caters to my own interests! They bring three threads together. First, they recall the long debate between using [normalized] Horvitz–Thompson and [self-normalized] Hájek estimators in survey sampling, pointing out that the latter is “usually the better estimator, despite estimation of an a priori known quantity” (citing from Särndal & al., 2003). Which is also my experience with importance sampling, as in this 1995 Note aux Comptes Rendus with George. The mathematical paradox of “estimating” a constant is central to other advances in the area, like noise-contrastive estimation à la Gutmann & Hyvärinen (2005) or the measure estimation of Kong & al. (2003). Even more interestingly, Datta & Polson consider there is a link with the inconsistent Bayesian (counter)example of Larry Wasserman and Jamie Robins, where the censoring probability increases with the value of the parameter of interest, paradox in which Chris Sims also got involved. (I was unaware that he had passed away last month.) And, lo and behold!, with the Stein “paradox” of my PhD years (and beyond).

Their central argument stands with the missing data link between Horvitz-Thompson survey sampling and Monte Carlo integration. (With a reference to our Riemann sum papers with Anne Philippe!, making me realise the authors had recently published an extension on that idea.) This reminds me very much of the missing measure approach of Kong et al. (2003). (Actually the reference appears in the final discussion.) The authors go over several paradoxes like Basu’s circus estimate (1988), Larry’s inconsistent Bayes estimate (2004), where Horvitz–Thompson performs nicely under compactness assumptions, the Bayesian answers (which include nested sampling even though I do not see the connection). Especially Li’s (2010) solution.  The attached numerical experiment displays a consistent underperformance of the Horvitz-Thompson estimator, in contrast with the theory… 

In conclusion, while enjoying very much revisiting so many examples and papers I came across in the past decades, I remain somewhat puzzled by the lack of overall message.

 

Bayesian sufficiency

Posted in Books, Kids, Statistics with tags , , , , , , , , , on February 12, 2021 by xi'an

“During the past seven decades, an astonishingly large amount of effort and ingenuity has gone into the search fpr resonable answers to this question.” D. Basu

Induced by a vaguely related question on X validated, I re-read Basu’s 1977 great JASA paper on the elimination of nuisance parameters. Besides the limitations of competing definitions of conditional, partial, marginal sufficiency for the parameter of interest,  Basu discusses various notions of Bayesian (partial) sufficiency.

“After a long journey through a forest of confusing ideas and examples, we seem to have lost our way.” D. Basu

Starting with Kolmogorov’s idea (published during WW II) to impose to all marginal posteriors on the parameter of interest θ to only depend on a statistic S(x). But having to hold for all priors cancels the notion as the statistic need be sufficient jointly for θ and σ, as shown by Hájek in the early 1960’s. Following this attempt, Raiffa and Schlaifer then introduced a more restricted class of priors, namely where nuisance and interest are a priori independent. In which case a conditional factorisation theorem is a sufficient (!) condition for this Q-sufficiency.  But not necessary as shown by the N(θ·σ, 1) counter-example (when σ=±1 and θ>0). [When the prior on σ is uniform, the absolute average is Q-sufficient but is this a positive feature?] This choice of prior separation is somewhat perplexing in that it does not hold under reparameterisation.

Basu ends up with three challenges, including the multinomial M(θ·σ,½(1-θ)·(1+σ),½(1+θ)·(1-σ)), with (n¹,n²,n³) as a minimal sufficient statistic. And the joint observation of an Exponential Exp(θ) translated by σ and of an Exponential Exp(σ) translated by -θ, where the prior on σ gets eliminated in the marginal on θ.