Archive for general public

probably overthinking it [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , on December 13, 2023 by xi'an

Probably overthinking it, written by Allen B. Downey (who wrote a series of books starting with Think, like Think Python, Think Bayes, Think Stats), belongs to this numerous collection of introductory books that aim at making statistics more palatable and enticing to the general public by making the fundamental concepts more intuitive and building upon real life examples. I would thus stop short of calling it “essential guide” as in the first flap of the dust jacket, since there exist many published books with a similar goal, some of which were actually reviews here. Now, there are ideas and examples therein I could borrow for my introductory stats course, except that I will cease teaching it next year! For instance, there are lots of examples related to COVID, which is great to engage (enrage?) the readers.

The book is quite pleasant to read, does not shy from mathematical formulae, and covers notions such as probability distributions, the Simpson, the Preston, the inspection, the Berkson paradoxes, and even some words on causality, sometimes at excessive lengths. (I have always been an adept of the concise church when it comes to textbook examples and fear that the multiplication of illustrations of a given concept may prove counterproductive.) The early chapters are heavily focussed on the Gaussian (or Normal) distribution. Making it appear as essential for conducting statistical analysis. When it does not, as in the ELO example, the explanations of a correction are less convincing.

I appreciated the book approach to model fit via the comparison of empirical cdfs with hypothetical ones. Also of primary interest is the systematic recourse to simulation, aka generative models, albeit without a systematic proper description. In the chapter (Chap 5) about durations, I think there are missed opportunities like the distributions of extremes (p 82) or the forgetfulness property of the Exponential distribution. Instead the focus is slightly diverging towards non-statistical issues on demography by the end of the chapter, with a potential for confusion between the Gomperz law and the Gomperz distribution. The Berkson paradox (Chap 6) is well-explained in terms of non-random populations (and reminded me when, years ago, when we tried to predict the first year success probability of undergrad applicants from their high school maths grade, the regression coefficient estimate ended up negative). Distributions of extremes do appear in Chap 8, if again seeking an ideal generic distribution seems to me rather misguided and misguiding. I would also argue that the author is missing the point of Taleb’s black swans by arguing in favour of a better modelling, when the later argues against the very predictability of extreme events in a non-stationary financial world… The chapter on fairness and fallacy (Chap 9) is actually about false positive/negative rates in different populations hence the ensuing unfairness (or the base fallacy). In that chapter there is no mention of Bayes (reserved for Think Bayes?!), but it is hitting hard enough at anti-vaxers (who will most likely not read the book). And does it again in the Simpson paradox chapter (Chap 10), whose proliferation is further stressed the following chapter on people becoming less racist or sexist or homophobic when they age, despite the proportion of racist/sexist/homophobic responses to a specific survey (GSS/Pew) increasing with age. This is prolonged into the rather minor final chapter.

Now that I have read the book, during a balmy afternoon in St Kilda (after an early start in the train to De Gaulle airport in freezing temperatures), I am a bit uncertain at what to make of it in terms of impact on the general public. For sure, the stories that accumulate chapter after chapter are nice and well argued, while introducing useful statistical concepts, but I do not see readers equipped enough to handle daily statistics with more than an healthy dose of scepticism, which obviously is a first step in the right direction!

Some nitpicking : the book is missing the historical connection to Quetelet’s “average man” when referring to the notion. And a potential explanation for the (approximate) log-Gaussianity of weights of individuals in a population through the fact that it is a volume, hence a third power of a sort.  Although birth weights are roughly Normal which kill my argument. I remain puzzled by the title, possibly missing a cultural reference (as there are tee-shirts sold with this sentence). It is the same as the name of a blog run by the author since 2011 and a fodder for the book. And the cover is terrible, breaking the words to fit the width making no sense, if I am not overthinking it! As often the book is rather US centric, although making no mention of US having much higher infant death rates than countries with similar GDPs when this data is discussed.

[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE.]

the joy of stats [book review]

Posted in Books, pictures, University life with tags , , , , , , , , , , , , on April 8, 2019 by xi'an

David Spiegelhalter‘s latest book, The Art of Statistics: How to Learn from Data, has made it to Nature Book Review main entry this week. Under the title “the joy of stats”,  written by Evelyn Lamb, a freelance math and science writer from Salt Lake City, Utah. (I noticed that the book made it to Amazon #1 bestseller, albeit in the Craps category!, which I am unsure is completely adequate!, especially since the book is not yet for sale on the US branch of Amazon!, and further Amazon #1 in the Probability and Statistics category in the UK.) I have not read the book yet and here are a few excerpts from the review, quoted verbatim:

“The book is part of a trend in statistics education towards emphasizing conceptual understanding rather than computational fluency. Statistics software can now perform a battery of tests and crunch any measure from large data sets in the blink of an eye. Thus, being able to compute the standard deviation of a sample the long way is seen as less essential than understanding how to design and interpret scientific studies with a rigorous eye.”

“…a main takeaway from the book is a sense of circumspection about our confidence in what is known. As Spiegelhalter writes, the point of statistical science is to ease us through the stages of extrapolation from a controlled study to an understanding of the real world, `and finally, with due humility, be able to say what we can and cannot learn from data’. That humility can be lacking when statistics are used in debates about contentious issues such as the costs and benefits of cancer screening.

a brief on naked statistics

Posted in Books, R, Statistics, University life with tags , , , , , , , , , on April 3, 2013 by xi'an

Over the last Sunday breakfast I went through Naked Statistics: Stripping the Dread from the Data. The first two pages managed to put me in a prejudiced mood for the rest of the book. To wit: the author starts with some math bashing (like, no one ever bothers to tell us about the uses of high school calculus!) either because he really feels like this or because it pays with the intended audience (like, we are on the same side, pal!), he then shows how he outsmarted his high school math teacher by spotting the exam was not possibly designed for his class and then another math teacher by just… re-inventing the steps leading to Zeno’s paradox (said Zeno of Elea not appearing in the credits of the book, to be sure) and sums it up with an NRA argument: “statistics is like a high-caliber weapon: helpful when used correctly” (p.xiv). Add to that a highly ethnocentric perspective that makes the book hardly readable for anyone outside the US, due to its absolute focus on all things American (exaggerating just a wee bit: who are Lebron James, Kim Kardashian, and Dan Rather?! what is Netflix?! why’s this Donald Rumsfeld guy quoted throughout the book?! how do they play baseball?! What do NBA, NHL, and SAT stand for?! &tc.)—as best illustrated by the facts that it took Charles Wheelan three months to realise a (golf) laser measuring instrument he had received could be in another unit that feet, namely meters!, and that he considers paying 100 rupees for a chai (मसाला चाय) in India a cheap price when this amount roughly corresponds to the average daily salary there…—. Top the whole thing with the fact that the author has already written a Naked Economics and seemingly found gold. (I am desperate for the incoming Naked  Paleopathology tome in the series!) And there you get me stuck with such a highly negative a priori about Naked Statistics that I could not shake it off for the rest of the book.

“This book will not make you a statistical expert (…) This book is not a textbook.” (p.xv)

With this warning in mind about my bias, let’s get on with what’s in this book. The above tells us what isn’t. To quote further from the author, the book “has been designed to introduce the statistical concepts with the most relevance to everyday life“ (p.xv). Naked Statistics goes over the basic notions of statistics (mean, standard deviation, correlation, linear regression, testing, design, polling), gives a sprinkle of probability background (counting models and the central limit theorem, which Wheelan considers as part of statistics), and spend the remaining chapters warning the reader(s) about the possible missuses of models and statistical tools if implemented in the wrong situations or with the wrong type of data. (There are a few graphs, but they are not particularly inspiring.) All this done with the minimum amount of maths formulae, mostly hidden in footnotes and appendices. (But then why adding an extra formula for σ when one is given just before for σ²?!) Sometimes, the minimum is not enough, as demonstrated by the “formula for calculating the correlation coefficient” (p.61) which takes a whole page of text to get around this absurdity of not using maths symbols like Σ and concludes with the lame “I’ll wave my hands and let the computer do the work” (p.61)! Somehow surprisingly, given the low-key nature of the book, it includes a final appendix on statistical software. From Excel, to SAS, Stata, and …R! While I am pleased at this inclusion, it sounds very much orthogonal to the purpose and the intended audience of Naked Statistics. I cannot fathom anyone reading the book and then immediately embarking upon writing an R code without stopping by a statistics textbook or formal training. (Incidentally, the author reproduces the usual confusion between free and open source, p.259.) Continue reading →