Archive for statistical literacy

Data science ethics [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on May 5, 2025 by xi'an

Data science ethics (concepts, techniques and cautionary tales), by David Martens, was published in 2022 by Oxford University Press. The book is inspired by the author’s  course on Data Science and ethics he has been teaching at the University of Antwerp. (With a link to his slides.) The 255p book proceeds by decomposing the ethics of data science into its different steps: data gathering (Chap. 2), data preprocessing (Chap. 3), modelling (Chap. 4), evaluation (Chap. 5), and deployment (Chap. 6). Following the `FAT Flow Framework´, where FAT stands for fairness, accountability, and transparency.

Do not expect much maths, stats, or anything quantitative: this book is mostly about concepts, even though some (mostly well-known) illustrations are provided. Chapter 2 includes a description of encryption (with homomorphic encryption treated in Chapter 4). And somewhat improbably quantum computing. Differential privacy gets a few pages with not a single formula (until Chapter 4, again).

Chapter 3 covers k-anonymity, record linkage, reidentification, (through cautionary tales) and discrimination through biases in the (learning) dataset.  Chapter 4 is defining ε differential privacy with the Laplace randomization as a possible implementation and with no critical stance on the limitations of the concept. The computation limitations of homomorphic encryption are more clearly pointed out. Federated learning is only quickly mentioned. The section about measuring fairness and reducing bias implies that some prior knowledge is available about whom is potentially discriminated and which covariates to add to the model. The last section on explicability of predictions is worthwhile in signalling the difficulty with most (black box) AI but the example opposing an SVM model to a logistic model is not tremendously convincing in that neither model is true.

Chapter 5 addresses the crucial challenge of ethical evaluation in a rather verbose and vague manner. For instance, with no instruction on how to resist adversarial attacks. Or criticising p-hacking and multiple testing while missing the elephant in the room (p-values!). Drifting from the topic when discussing the misdeeds of Diederik Stapel. Most of the same goes about Chapter 6 and its take on ethical deployment, when going through examples such as Google’s policies in China. Or general musing on the impact of AI on societal inequalities (with mentions of companies and CEOs who have since then back-pedalled on their ethical engagement). These chapters are lacking in tools and (more) practical recommendations.

One interesting aspect of the book is the attention paid to the EU(ropean) aspect of these concerns, through the GDPR (General DAta Protection Regulations) adopted by the European Parliament in 2016. (There is also a brief mention of China’s regulations, but no details beyond a reference. Maybe the Chinese edition differs.)

the polls weren’t wrong [alt book review]

Posted in Statistics with tags , , , , , , , , , , , , , , , on April 1, 2025 by xi'an

the polls weren’t wrong [book review]

Posted in Books, R, Statistics, University life with tags , , , , , , , , , , , , , , , , on November 1, 2024 by xi'an

While Nate Silver and his colleagues (as The New York Times Nate Cohn) have brought (part of) the general public to adopt a (more) scientific perspective on political polls, the way to combine them, and the need to keep uncertainty fully quantified (witness this recent “Two Theories for Why the Polls Failed in 2020, and What It Means for 2024” by Nate Cohn for the NYT’s Tilt), the author of this book embarks upon a crusade (and a lengthy rant) against pollsters and analysts and media reporters, with the single, many times repeated, argument that non-responses and undecided voters are crucial for the final election outcome… And that a poll gives a snapshot of the current (time) population opinion, not a prediction of its future state. There is not the slightest trace of statistical depth, there is actually no statistics at all found throughout the book, apart from a section (p.281) entitled “Threats to inferential statistics” (with not no maths either, as a few ratio manipulations and a chapter title invoking Jakob Bernoulli!) do not count as maths!, but a lot of repetitions on the same theme and dismissal of statisticians’ analyses, like Nate Silver’s. Opposing to them the theories of Nick Panagakis, a 1990’s pollster. (Funny enough, the Amazon reviews include one “expert in inferential statistics, the major tool employed by Carl [Alen] in this book” and another one stating that “Carl Allen takes the reader through a journey towards statistical literacy“!) And an R code (p.52) for plotting the outcome of 30 Binomial random draws.

“Unless my efforts achieve far more notoriety than even my most optimistic forecast would predict, [the Proportional Method] is unlikely to go away any time soon.” p.183

“If it seems like I’m picking on FiveThirtyEight a lot, it’s not because there are no other forecasters who are better or worse.” p.235

“Sounding eerily like myself, [Nate Silver] pointed out [in 2008] that `many things can happen’ months before the election.” p.208

Reading through the book (during a trip to & from Warwick) was painful, both for the feeling of being stuck in a plane with a perfect unknown, next seat, trying to force their weird theory upon you and no way to escape their rant, as well as for the terrible style of said book, full of repetitions and one-sentence paragraphs. (Of course, nothing as bad as this time near the 2012 US elections I flew to Des Moines next to an inebriated woman that would not stop blathering about her life!) Or as I imagine a card game addict defending their martingale as a sure way to win against the casino. It is also the first time I see references repeated (in postcripts) as many times as they are cited within a chapter. There is no true insight on how polling companies construct their polling samples, how they post-process outcomes by regression techniques, and no reflection on the unique weirdness of the US electoral system in that a few States determine the outcome (rather than majority votes) and thus how a tiny number of voters (escaping the law of Large Numbers) hold the overall result in their hand.

Thus (as most readers will have forecasted) concluding by not recommending the book!

[Disclaimer about potential self-plagiarism: this post or an edited version of it could possibly appear in my Books Review section in CHANCE. Most unlikely though!]