Archive for fairness

Congrats, IOC: discriminatory sex testing, but inclusion of Russian athletes!

Posted in Kids, Running with tags , , , , , , , , , , , , , , , on April 6, 2026 by xi'an

persuasive (and Oceanic) privacy

Posted in Books, Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on February 3, 2026 by xi'an

I am quite excited about the paper James Baillie, Joshua Bon, Judith Rousseau, and myself just arXived! A novel framework for measuring privacy we have been working on for at least the past year, partly through the previous Les Houches privacy workshops. In the spirit of these workshops and the larger scale ERC Synergy grant OCEAN, we develop therein a rather generic Bayesian game-theoretic perspective on achieving statistical privacy. It involves a Sender (observing the original data and delivering a limited output) and a Receiver (with potential adversarial intentions). The paper mostly focus on setting a theoretical framework, including the creation of new, purpose-driven privacy definitions that are rigorously justified, while also allowing for the assessment of existing privacy guarantees through game theory. While this was not our original intent, we show that pure and probabilistic differential privacy notions, in the Dwork et al. (2006) sense, are special cases of our framework. This setting provides new interpretations of the post-processing inequality. Furthermore, and somewhat more importantly, we also prove that our privacy guarantees can be established for deterministic algorithms, which are outside current privacy standards. Hopefully, we’ll make further progress at the incoming privacy workshop next month, to be held in Venice (again).

Data science ethics [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on May 5, 2025 by xi'an

Data science ethics (concepts, techniques and cautionary tales), by David Martens, was published in 2022 by Oxford University Press. The book is inspired by the author’s  course on Data Science and ethics he has been teaching at the University of Antwerp. (With a link to his slides.) The 255p book proceeds by decomposing the ethics of data science into its different steps: data gathering (Chap. 2), data preprocessing (Chap. 3), modelling (Chap. 4), evaluation (Chap. 5), and deployment (Chap. 6). Following the `FAT Flow Framework´, where FAT stands for fairness, accountability, and transparency.

Do not expect much maths, stats, or anything quantitative: this book is mostly about concepts, even though some (mostly well-known) illustrations are provided. Chapter 2 includes a description of encryption (with homomorphic encryption treated in Chapter 4). And somewhat improbably quantum computing. Differential privacy gets a few pages with not a single formula (until Chapter 4, again).

Chapter 3 covers k-anonymity, record linkage, reidentification, (through cautionary tales) and discrimination through biases in the (learning) dataset.  Chapter 4 is defining ε differential privacy with the Laplace randomization as a possible implementation and with no critical stance on the limitations of the concept. The computation limitations of homomorphic encryption are more clearly pointed out. Federated learning is only quickly mentioned. The section about measuring fairness and reducing bias implies that some prior knowledge is available about whom is potentially discriminated and which covariates to add to the model. The last section on explicability of predictions is worthwhile in signalling the difficulty with most (black box) AI but the example opposing an SVM model to a logistic model is not tremendously convincing in that neither model is true.

Chapter 5 addresses the crucial challenge of ethical evaluation in a rather verbose and vague manner. For instance, with no instruction on how to resist adversarial attacks. Or criticising p-hacking and multiple testing while missing the elephant in the room (p-values!). Drifting from the topic when discussing the misdeeds of Diederik Stapel. Most of the same goes about Chapter 6 and its take on ethical deployment, when going through examples such as Google’s policies in China. Or general musing on the impact of AI on societal inequalities (with mentions of companies and CEOs who have since then back-pedalled on their ethical engagement). These chapters are lacking in tools and (more) practical recommendations.

One interesting aspect of the book is the attention paid to the EU(ropean) aspect of these concerns, through the GDPR (General DAta Protection Regulations) adopted by the European Parliament in 2016. (There is also a brief mention of China’s regulations, but no details beyond a reference. Maybe the Chinese edition differs.)

more oceanographers in Les Houches

Posted in Books, Kids, Mountains, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , on February 2, 2025 by xi'an

ridge6

Our second internal research workshop on privacy, covered  by our ERC Synergy project OCEAN is taking place in Les Houches, French Alps, this very week with 12 researchers gathering for further brain-storming on some of the themes at the core of the project, like algorithmic tools for multiple decision-making agents, along with Bayesian uncertainty quantification and Bayesian learning under constraints (scarcity, fairness, privacy). Due to the small size of the workshop, it will take place in the same local (and uncommon) hotel as last year. And where I had a genuine blackboard delivered (to be donated afterwards to École de Physique des Houches), as the place is not used to mathematicians discussing on and around a blackboard. Even though the schedule does not leave much time for skiing, I hope for a few opportunities, given that the ski cover is quite good this year!

probably overthinking it [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , on December 13, 2023 by xi'an

Probably overthinking it, written by Allen B. Downey (who wrote a series of books starting with Think, like Think Python, Think Bayes, Think Stats), belongs to this numerous collection of introductory books that aim at making statistics more palatable and enticing to the general public by making the fundamental concepts more intuitive and building upon real life examples. I would thus stop short of calling it “essential guide” as in the first flap of the dust jacket, since there exist many published books with a similar goal, some of which were actually reviews here. Now, there are ideas and examples therein I could borrow for my introductory stats course, except that I will cease teaching it next year! For instance, there are lots of examples related to COVID, which is great to engage (enrage?) the readers.

The book is quite pleasant to read, does not shy from mathematical formulae, and covers notions such as probability distributions, the Simpson, the Preston, the inspection, the Berkson paradoxes, and even some words on causality, sometimes at excessive lengths. (I have always been an adept of the concise church when it comes to textbook examples and fear that the multiplication of illustrations of a given concept may prove counterproductive.) The early chapters are heavily focussed on the Gaussian (or Normal) distribution. Making it appear as essential for conducting statistical analysis. When it does not, as in the ELO example, the explanations of a correction are less convincing.

I appreciated the book approach to model fit via the comparison of empirical cdfs with hypothetical ones. Also of primary interest is the systematic recourse to simulation, aka generative models, albeit without a systematic proper description. In the chapter (Chap 5) about durations, I think there are missed opportunities like the distributions of extremes (p 82) or the forgetfulness property of the Exponential distribution. Instead the focus is slightly diverging towards non-statistical issues on demography by the end of the chapter, with a potential for confusion between the Gomperz law and the Gomperz distribution. The Berkson paradox (Chap 6) is well-explained in terms of non-random populations (and reminded me when, years ago, when we tried to predict the first year success probability of undergrad applicants from their high school maths grade, the regression coefficient estimate ended up negative). Distributions of extremes do appear in Chap 8, if again seeking an ideal generic distribution seems to me rather misguided and misguiding. I would also argue that the author is missing the point of Taleb’s black swans by arguing in favour of a better modelling, when the later argues against the very predictability of extreme events in a non-stationary financial world… The chapter on fairness and fallacy (Chap 9) is actually about false positive/negative rates in different populations hence the ensuing unfairness (or the base fallacy). In that chapter there is no mention of Bayes (reserved for Think Bayes?!), but it is hitting hard enough at anti-vaxers (who will most likely not read the book). And does it again in the Simpson paradox chapter (Chap 10), whose proliferation is further stressed the following chapter on people becoming less racist or sexist or homophobic when they age, despite the proportion of racist/sexist/homophobic responses to a specific survey (GSS/Pew) increasing with age. This is prolonged into the rather minor final chapter.

Now that I have read the book, during a balmy afternoon in St Kilda (after an early start in the train to De Gaulle airport in freezing temperatures), I am a bit uncertain at what to make of it in terms of impact on the general public. For sure, the stories that accumulate chapter after chapter are nice and well argued, while introducing useful statistical concepts, but I do not see readers equipped enough to handle daily statistics with more than an healthy dose of scepticism, which obviously is a first step in the right direction!

Some nitpicking : the book is missing the historical connection to Quetelet’s “average man” when referring to the notion. And a potential explanation for the (approximate) log-Gaussianity of weights of individuals in a population through the fact that it is a volume, hence a third power of a sort.  Although birth weights are roughly Normal which kill my argument. I remain puzzled by the title, possibly missing a cultural reference (as there are tee-shirts sold with this sentence). It is the same as the name of a blog run by the author since 2011 and a fodder for the book. And the cover is terrible, breaking the words to fit the width making no sense, if I am not overthinking it! As often the book is rather US centric, although making no mention of US having much higher infant death rates than countries with similar GDPs when this data is discussed.

[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE.]