Archive for survey sampling

inverse probability weighting

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , on May 4, 2026 by xi'an

Quite recently, Jyotishka Datta and Nick Polson published a fairly interesting [imho] paper in The New England Journal of Statistics in Data Science, entitled Inverse Probability Weighting: From Survey Sampling to Evidence Estimation that (obviously) caters to my own interests! They bring three threads together. First, they recall the long debate between using [normalized] Horvitz–Thompson and [self-normalized] Hájek estimators in survey sampling, pointing out that the latter is “usually the better estimator, despite estimation of an a priori known quantity” (citing from Särndal & al., 2003). Which is also my experience with importance sampling, as in this 1995 Note aux Comptes Rendus with George. The mathematical paradox of “estimating” a constant is central to other advances in the area, like noise-contrastive estimation à la Gutmann & Hyvärinen (2005) or the measure estimation of Kong & al. (2003). Even more interestingly, Datta & Polson consider there is a link with the inconsistent Bayesian (counter)example of Larry Wasserman and Jamie Robins, where the censoring probability increases with the value of the parameter of interest, paradox in which Chris Sims also got involved. (I was unaware that he had passed away last month.) And, lo and behold!, with the Stein “paradox” of my PhD years (and beyond).

Their central argument stands with the missing data link between Horvitz-Thompson survey sampling and Monte Carlo integration. (With a reference to our Riemann sum papers with Anne Philippe!, making me realise the authors had recently published an extension on that idea.) This reminds me very much of the missing measure approach of Kong et al. (2003). (Actually the reference appears in the final discussion.) The authors go over several paradoxes like Basu’s circus estimate (1988), Larry’s inconsistent Bayes estimate (2004), where Horvitz–Thompson performs nicely under compactness assumptions, the Bayesian answers (which include nested sampling even though I do not see the connection). Especially Li’s (2010) solution.  The attached numerical experiment displays a consistent underperformance of the Horvitz-Thompson estimator, in contrast with the theory… 

In conclusion, while enjoying very much revisiting so many examples and papers I came across in the past decades, I remain somewhat puzzled by the lack of overall message.

 

Measuring abundance [book review]

Posted in Books, Statistics with tags , , , , , , , , , , , , on January 27, 2022 by xi'an

This 2020 book, Measuring Abundance:  Methods for the Estimation of Population Size and Species Richness was written by Graham Upton, retired professor of applied statistics, for the Data in the Wild series published by Pelagic Publishing, a publishing company based in Exeter.

“Measuring the abundance of individuals and the diversity of species are core components of most ecological research projects and conservation monitoring. This book brings together in one place, for the first time, the methods used to estimate the abundance of individuals in nature.”

Its purpose is to provide a collection of statistical methods for measuring animal abundance or lack thereof. There are four parts: a primer on statistical methods, going no further than maximum likelihood estimation and bootstrap. The term Bayesian only occurs once, in connection with the (a-Bayesian) BIC. (I first spotted a second entry, until I realised this was not a typo and the example truly was about Bawean warty pigs!) The second part is about stationary (or static) individuals, such as trees, and it mostly exposes different recognised ways of sampling, with a focus on minimising the surveyor’s effort. Examples include forestry sampling (with a chainsaw method!) and underwater sampling. There is very little statistics involved in this part apart from the rare appearance of a MLE with an asymptotic confidence interval. There is also very little about misspecified models, except for the occasional warning that the estimates may prove completely wrong. The third part is about mobile individuals, with capture-recapture methods receiving the lion’s share (!). No lion was actually involved in the studies used as examples (but there were grizzly bears from Yellowstone and Banff National Parks). Given the huge variety of capture-recapture models, very little input is found within the book as the practical aspects are delegated to R software like the RMark and mra packages. Very little is written on using covariates or spatial features in such models, mostly dedicated to printed output from R packages with AIC as the sole standard for comparing models. I did not know of distance methods (Chapter 8), which are less invasive counting methods. They however seem to rely on a particular model of missing on individuals as the distance increases. The last section is about estimating the number of species. With again a model assumption that may prove wrong. With the inclusion of diversity measures,

The contents of the book are really down to earth and intended for field data gatherers. For instance, “drive slowly and steadily at 20 mph with headlights and hazard lights on ” (p.91) or “Before starting to record, allow fish time to acclimatize to the presence of divers” (p.91). It is unclear to me how useful the book would prove to be for general statisticians, apart from revealing the huge diversity of methods actually employed in the field. To either build upon these or expose students to their reassessment. More advanced books are McCrea and Morgan (2014), Buckland et al. (2016) and the most recent Seber and Schofield (2019).

[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Book Review section in CHANCE.]

Cédric Villani on COVID-19 [and Zoom for the local COVID-19 seminar]

Posted in Statistics, University life with tags , , , , , , , , , , , , on June 19, 2020 by xi'an

From the “start” of the COVID-19 crisis in France (or more accurately after lockdown on March 13), the math department at Paris-Dauphine has run an internal webinar around this crisis, not solely focusing on the math or stats aspects but also involving speakers from other domains, from epidemiology to sociology, to economics. The speaker today was [Field medalist then elected member of Parliament] Cédric Villani, as a member of the French Parliament sciences and technology committee, l’Office parlementaire d’évaluation des choix scientifiques et technologiques (OPECST), which adds its recommendations to those of the several committees advising the French government. The discussion was interesting as an insight on the political processing of the crisis and the difficulties caused by the heavy-handed French bureaucracy, which still required to fill form A-3-b6 in emergency situations. And the huge delays in launching a genuine survey of the range and diffusion of the epidemic. Which, as far as I understand, has not yet started….

Estudio nacional epidemiológico de la infección por SARS-CoV2 en España [a proper survey]

Posted in Mountains, pictures, Statistics, Travel with tags , , , , on April 28, 2020 by xi'an

[A proper survey on the prevalence of the virus in the Spanish population has been launched, with 36,000 representative households chosen from census bases. Thanks to Victor for pointing out the survey!]

  • Se desarrollará en las próximas semanas en colaboración con los servicios de salud de las CCAA.

  • La Atención Primaria tendrá un papel relevante en la realización de un estudio que pretende estimar el porcentaje de la población española que ha desarrollado anticuerpos frente al nuevo coronavirus.

  • En colaboración con el INE, se han seleccionado más de 36.000 hogares españoles, para que la muestra tenga participantes de todos los grupos de edad y localizaciones geográficas.

27 de abril de 2020.- Hoy comienza el Estudio Nacional Epidemiológico de la infección por SARS-CoV2 en España (ENE-COVID), diseñado por el Ministerio de Sanidad y el Instituto de Salud Carlos III (ISCIII) con la colaboración de las CCAA.

Las CCAA proporcionarán el personal sanitario para la realización del proyecto y serán las encargadas de adecuar la logística del estudio de la forma que se considere más adecuada en cada territorio, garantizando que se cumplen todos los requisitos metodológicos del estudio.

A través de las Consejerías de Sanidad o de los propios centros de salud, se irá citando a los participantes para la obtención de muestras. Las llamadas comenzarán este mismo lunes. “La participación es totalmente voluntaria”, ha destacado el ministro de Sanidad, Salvador Illa, “pero aprovecho para animar a todas las personas que sean contactadas a participar en el estudio. Los resultados serán de enorme utilidad para toda la sociedad española.

Con este estudio, el Ministerio de Sanidad y el ISCIII, dependiente del Ministerio de Ciencia e Innovación, en estrecha colaboración con las Comunidades Autónomas, pretenden estimar el porcentaje de la población española que ha desarrollado anticuerpos frente al nuevo coronavirus SARSCoV-2 (concepto conocido como seroprevalencia). La información obtenida será de enorme relevancia para la toma de decisiones de salud pública en el conjunto del Estado.

El papel de los servicios de Atención Primaria de Salud será especialmente relevante a lo largo de todo el proceso.

another surmortality graph

Posted in Books, R, Statistics with tags , , , , , , , , on April 27, 2020 by xi'an

Another graph showing the recent peak in daily deaths throughout France as recorded by INSEE and plotted by Baptiste Coulmont from Paris 8 Sociology Department. And further discussed by Arthur Charpentier on Freakonometrics. With a few days off due to reporting, this brings an objective perspective on the impact of the epidemics (and of the quarantine) compared with the other years since 2001, without requiring tests or even surveys. (The huge peak in August 2003 was an heat wave that decimated elderly citizens throughout France.)