Archive for capture-recapture

independent Gaza Mortality Survey reports more than 80,000 fatalities

Posted in Books, Kids, Statistics with tags , , , , , , , , , , , on July 13, 2025 by xi'an

Nature of 27 June reported on an independent survey that interviewed 2000 households in the Gaza strip, except in the most dangerous areas, in December 2024. As noted by Patrick Ball, director of research at Human Rights Data Analysis Group (who took part in Datascience for Good at CIRM), it is amazing a proper survey could be conducted in such dramatic conditions, with consistent findings given an earlier study using capture-recapture published in The Lancet, and exceeds the official figures published by the Ministry of Health in Gaza. In part because they included non-violent deaths since October 2023.  As stressed by the authors, assessing the magnitude of the death numbers (4% of the Gaza population) and of the unique proportion of non-combatant deaths (over 60%) in a modern conflict, but numbers do not replace a

“…project of memorialization [that] has only begun and must continue for many years after the war ends. Estimates of numbers killed, the focus of the present paper, can help guide this casualty recording work. But we cannot provide each human being due recognition in death simply by estimating the number of deaths.”

A modern introduction to probability and statistics [book review]

Posted in Books, R, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , on July 12, 2025 by xi'an

In the plane to Bengaluru, I read through the book A modern introduction to probability and statistics, by Graham Upton—whose Measuring Animal Abundance I reviewed for CHANCE a while ago—, which is based on the earlier Understanding Statistics, written jointly with Ian Cook. (Not to be confused with A modern introduction to probability and statistics by Dekking et al.) The subtitle is understanding statistical principles in the computer age. Sorry, in the age of the computer. While the cover is most pleasant (and modern), as noticed by an AF flight attendant, the contents are very very standard and could have been written decades ago since the main concession to “the” computer age is the inclusion of a few R commands at the end of most chapters. There are even a few distribution tables here and there (in case “the” computer is not available). But there is no other connection with computational statistics or statistical computing.

The classicism of the contents and the intended audience mean there is little therein on which to either object or criticise. The mixture of elementary probability and basic statistics in a single textbook always feels awkward to me and I think I would have trouble teaching solely from this material. Apart from the glaring typo on the variance of the sum of two correlated random variables on page 87, missing the factor 2 in front of the covariance, while correct(ed) p97 (and the inevitable “the the” typo spotted once). My main criticisms are on the potential confusion between samples and populations in the early chapters, when some statistics are used as motivational examples, as for instance in a (hidden) Monte Carlo stabilisation to the limiting values (p57), way before the Law of Large Numbers is introduced,, the variable mileage in mathematical rigour (while being uncertain that first year students can handle integrals and derivatives), the textbook examples, and the amount of the book contents spent on descriptive statistics and even more on the “classical” tests, with no critical perspective on using point nulls or p-values. The book concludes with a four page (benevolent) chapter on Bayesian statistics that is superfluous imho, or even counterproductive since my experience with a rushed introduction to Bayesian principles almost always result in a rejection of said principles. Plus, the illustration with the coin tossing is not particularly helpful since Andrew maintains that one can load a die, but cannot bias a coin. (A similar reservation on the half-page 289 coverage on pseudo-random generation and Monte Carlo principles for computing p-values.)

Minor (mostly idiosyncratic) remarks follow: CLT prior to LLN,   n-1 in sample sd, little to no model criticism (ntbcf goodness of fit), missing an opportunity when mentioning the varying probability of a day being a birthday (p31) in contrast with BDA cover story, and another opportunity to cite the 2024 Ig Nobel Prize for coin tossing around the LLN, an unclear definition for random variables( p53) and a potentially confusing introduction of Poisson distributions through a informal reference to Poisson processes (and no reason why the years of accession of the kings of Sussex and England till Guillaume—making a return on p178 with the Domesday Book—in 1066 should follow such a process as suggested in Figure 3.5), a surprising definition of the constant e as the special case of exp(x) when x=1 and its series expansion (p70), omitting proofs on laws of sums of iid rv’s by introducing moment generating functions rather late, another obscure reference to a 16th German treatise on surveying as a precursor of the CLT (p131), a proof for the normalising constant of the Normal density that will most likely escape most first year students, a introduction of the t, F, and χ² distributions with no mention of their respective densities (pp141-147), never defining a joint Normal distribution density, insisting on unbiasedness without noting that maximum likelihood—with a strange motivation that it “makes the next sample of n observations most likely to resemble the data in the current sample (p228)—estimators are almost always biased, an abundance of footnotes that may prove of little interest for the youngest readers.

[Disclaimer about potential self-plagiarism as usual: this post or an edited version will eventually appear in my Books Review section in CHANCE.]

Bayesian Inference: Theory, Methods, Computations [book review]

Posted in Statistics with tags , , , , , , , , , , , , , , , , , , , , , on November 12, 2024 by xi'an

Bayesian Inference: Theory, Methods, Computations by Silvelyn Zwanzig and Rauf Ahmad, both from Uppsala University, is a recent book published by Chapman & Hall / CRC Press. About 300p long (plus appendices), it covers the core aspects of Bayesian inference, namely the decision theoretic motivations, its asymptotic validation, the specifics of estimation and testing, and the computational approximations (MC, MCMC, ABC, VB), with entries on prior specification and Normal linear models. And some R codes. It is (and feels like) constructed from Master and PhD courses (at Uppsala University), with a rigorous mathematical presentation and many examples, some related to biostatistics. Drawings from the first author’s daughter are included in most chapters, to this reviewer’s bemusement. From a further personal viewpoint, the book also reads rather close to my (Bayesian) choice of a Bayesian textbook, which proves rather accurate since several chapters are inspired by my own Bayesian Choice. as acknowledged therein. As well as by the more recent Statistical Decision Theory: Estimation, Testing, and Selection by Liese & Miescke (2008) and Introduction to the Theory of Statistical Inference by Liero & Zwanzig (2011). Witness, for instance, an example of prior construction for capture-recapture experiments on lizards as analysed by my PhD student Dupuis (1995) [with a curious switch to the authors on p.263] and  also included in The Bayesian Choice (with drawing 2.9 incorrect in that the lizards there have marks on their backs, instead of the code adopted by the ecologists, namely cutting one specific phalange for each capture).

Other minor quandaries: The usual issue of quoting the wrong edition for creating a method, as when citing Jeffreys (1946) for inventing non-informative priors [p.53], failing to point out the parameterisation invariance of intrinsic losses [p.95]considering that Bayes factors are only relevant for obtaining evidence against the null hypothesis [p.216], recommending BIC and DIC (!) [pp.232-6], advocating sampling importance resampling (SIR) for approximate sampling from the target (omitting infinite variance issues) [p.253], defining annealing as using “several trial distributions” [p.261], a mistake in ABC-MCMC [p.274] since the case when the simulated data is too far from the actual data should lead to a repetition rather than a pure rejection.

All in all, a reasonable textbook with some recent input, but still lacking in originality, if I may subjectively say so.

[Disclaimer about potential self-plagiarism: this post or an edited version of it could possibly appear in my Books Review section in CHANCE.]

how many migrants died trying to reach Europe?

Posted in Books, Kids, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , on February 6, 2024 by xi'an

“We estimated that about 40,000 human beings have died trying to enter the European Union, during about 5500 tragic attempts, in the period between January 1993 and March 2019.”

A paper by Alessio Farcomeni published last year in Annals of Applied Statistics [that is delivered to me by regular mail, but which I missed] about estimating the number of deaths on migration routes. I have been interested in this question for moral and statistical reasons since the beginning of the Syrian civil war and the induced massive increase in the number of people attempting to cross the Mediterranean Sea to reach Europe. Unfortunately, despite different attempts to contact governmental agencies (like the French Minister of the Interior),  NGOs (like Amnesty), friends, academics (incl. a Dutch group on that very topic but with no data or data scientist), and connections (with the Italian Navy, the Tunisian government and Frontex), I could not access more than newspaper level data or highly local data that did not allow for a general picture.


“The law of large numbers guarantees consistency as long as the population size estimator is consistent and the model is well-specified.”

As it was my intention, the paper uses capture-recapture. The data is obtained from UNITED for Intercultural Action. It had recorded 4333 attempts to enter Europe that had at least one death occurring, between January 1993 and March 2019, as reported by one or several sources, like newspaper articles. Sources that are most often not independent, but rather copying one another, which is obviously problematic for the extrapolation to the unreported cases and for resorting to capture-recapture estimation. And the number of deaths per event is itself most likely inexact, since casualties are not always recovered or identified as such. Since migration routes, migrant flows, and smuggler policies keep massively changing, the homogeneity of the observations over nearly 30 years is low to inexistent, which makes invoking consistency rather inappropriate. This is also the reason why I find the approach followed by the paper too strongly model-based as for instance when relying on an Horvitz-Thompson estimator or using a GLM to link the number of deaths in one crossing and the number of sources reporting the tragedy. This fundamental difficulty in modelling or inferring from such unreliable and untrustworthy data sources and the absence of record linkage with other datasets like the entries in the border countries (e.g., Turkey) or the number of prevented crossings by the local coastguards alas make the final estimate of 40,000 deaths at sea close to impossible to calibrate from a model-free perspective. The actual figure is not only higher, but maybe considerably so. Unfortunately, the datasets that would allow linkage and recapture are unavailable or inexistent in departure countries, while arrival countries most abstain from storing data about the migrant flows and histories…

capture & maybe recapture

Posted in Books, pictures, Running, Statistics, University life with tags , , , , , , , , , on September 5, 2023 by xi'an

I read population size estimation with capture-recapture in presence of individual misidentification and low recapture arXived by Rémy Fraysse and coauthors on my flight back from Saigon. The setup is one of a capture-recapture experience where potential misidentification (of a recapture individual labelled as new) may occur due to visual identification errors as, e.g., in whale studies. Trying to handle the issue, Yoshizaki et al. (2011) proposed adding one layer to the temporal Darroch model M[t] via a probability α of creating a “ghost” (by failing to recognise a formerly observed individual). When representing the experiment as a partly observed Markov process (as in Dupuis, 1995), this addition brings another completely latent process for misidentification. Completely in the sense that a misidentification is never observed (while a proper identification is). This means there is an issue with identifiability between failing to capture and misidentification, as a new individual may be captured for the first time or misidentified on that round.

Processing this model (i.e., producing a simulation algorithm of the posterior) can be done formally by a Gibbs completion (as in Dupuis, 1995) but this may prove a non-irreducible scheme (Schofer & Bonner, 2015), a problem solved by considering instead Metropolis-Hastings steps. The current paper is an extension of the above to the multiple states Arnason-Schwarz model with no theoretical convergence issue, besides running a large sized completion. It is mostly a simulation experiment with a comparison on different priors on the misspecification rate, some highly informative, others not, with a frequentist assessment of coverage. Given the identifiability issue mentioned above, this is not particularly helpful since it simply exhibits the fact that with the right priors the parameter values are unbiasedly estimated, while a low recapture plus high misidentification setting makes estimation more difficult.