Archive for biases
the thought police
Posted in Statistics with tags 1984, baby Trump, biases, diversity, equity, ethnicity, George Orwell, homophobia, LGBT, LGBT rights, Native Americans, Newspeak, Ninety-Eighty-Four, NYT, race, racial profiling, stereotypes, The New York Times, Tought Police, transphobia, Trump administration on March 31, 2025 by xi'anHands-On Differential Privacy [book review]
Posted in Books, R, Statistics, University life with tags Bayesian inference, biases, book review, CHANCE, classification, coding, contextual integrity, differential privacy, GANs, Harvard, Laplace distribution, LaTeX, Lipschitz continuity, machine learning, O'Reilly Media, OpenDP, Python, slate pencil sea urchin, Type I error, Type II error on October 2, 2024 by xi'anHands-On Differential Privacy was published just a few months ago (from September 2024!) by (the US publisher) O’Reilly, famous for its programming and technical books with animal covers! A slate pencil sea urchin in the present case. The book is indeed classical O’Reilly’s, with lots of notes, little theory (or maths!) and symbols, a loose structuring of the chapters (no section numbers) and highly detailed examples, and of course plenty of OpenDP code inserts. For instance, in the present case, a case study about the privatization of a sample average x̄ that takes about ten pages. Terrible equation rendering btw (what’s wrong with LATEX?!). Overall, I am quickly lost in most of the chapters due to a lack of a driving narrative, facing instead a catalogue of possible scenari and procedures, appearing one after the other as in a fashion show.
Hands-On Differential Privacy is written by Ethan Cowan, Michael Shoemate, and Mayana Pereira. I came across the book during the OpenDP workshop at Harvard [that took place right after my return from the Pacific Northwest] and it is definitely linked with OpenDP, all authors being actually involved at one stage or another in the OpenDP Team. The style of the book is once again in tune with the O’Reilly manuals, which sort of clashes with my preferences. For instance, the introduction of differential privacy (Chapter 2) is quite extensive. Chapter 3 proceeds to teach about private data transform(ation)s, stability (a rewording of Lipschitz-ianity), with code illustrations, often repeating the earlier derivation (see eg p203), while Chapter 4 is its equivalent for private mechanisms. (With the diagrams Figures 3-1 and 4-1 differing only in highlighting/bolding different functions in a privatized data processing pipeline.) Returning to differential privacy with a privacy loss parameter and to Laplace and exponential mechanisms, Chapter 5 proposes several notions of privacy, all closed under post-processing. This includes Wasserman and Zhou (2010) interpretation of privacy as hypothesis testing, except it is not exploited further than connecting type I and type II with (ε,δ) parameters. Chapter 6 concludes Part I about concepts with a series of (fearless) combinators, keeping stability and privacy. With an increasing proportion of coding excerpts which I [imho] did not find particularly helpful.
Nothing about statistical loss of information or efficiency, bias, &tc. until Chapter 8 (p199) and even then so little. Part II is about practice, with a first Chapter 7 on setting a privacy unit (e.g., a person-month) before ensuring their privacy is protected. And discussing unbounded contributions (not unbounded data!). While Chapter 8 very thinly covers statistical modelling, while remaining agnostic about the choice of statistical procedures (Bayes being solely and naïvely mentioned for classification, furthermore with data-based evaluation of the class “prior” probabilities, p211). At this stage, procedures are often only defined through spinets of code, like the private Theil-Sen estimator (pp204-205). The continuous case boils to a Normality assumption, with its pmf being defined (p212) as
which contains at least three errors! Chapter 9 is the equivalent of Chapter 8 for machine learning, mostly centred on private gradient descent. And a Pytorch section (pp232-235). Completed by a light Chapter 10 on synthetic data, which does not seem to broach upon the issue of large dimension covariates, providing instead a list of GAN synthetizers.
Part III (Deploying differential privacy) is even more about practice, with Chapter 11 on privacy attacks, Chapter 12 on calibrating a privacy mechanism (co-written with Jayshree Sarathy), and good practice (like codebooks and data annotations), with the appearance of contextual integrity I discovered if not perfectly understood last year at the BIRS workshop in Kelowna. And Chapter 13 on planning a privacy project, with an 11 step checklist, most of which are quite vague [imho] and do include strategies to make the data owners confident their privacy is safe.
[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE]
Nature worries
Posted in Statistics with tags Amazon, biases, Brazil, Brexit, China, EU, facial recognition, Google, India, Italian politics, John Ioannidis, Nature, self-citations on October 9, 2019 by xi'anIn the 29 August issue, worries about the collapse of the Italian government coalition for research (as the said government had pledge to return funding to 2009 [higher!] levels), Brexit as in every issue (and the mind of every EU researcher working in the UK), facial recognition technology that grows much faster than the legal protections which should come with it, thanks to big tech companies like Amazon and Google. In Western countries, not China… One acute point in the tribune being the lack of open source software to check for biases. More worries about Amazon, the real one!, with Bolsonaro turning his indifference if not active support of the widespread forest fires into a nationalist campaign. And cutting 80,000 science scholarships. Worries on the ethnic biases in genetic studies and the All of Us study‘s attempt to correct that (study run by a genetic company called Color, which purpose is to broaden the access to genetic counseling to under-represented populations). Further worries on fighting self-citations (with John Ioannidis involved in the analysis). With examples reaching a 94% rate for India’s most cited researcher.
Statistics and Health Care Fraud & Measuring Crime [ASA book reviews]
Posted in Books, Statistics with tags American Statistical Association, ASA, biases, book review, CHANCE, Chicago, CRC Press, crime, crime statistics, David Spiegelhalter, Edith Abbott, health care, health care fraud, Minority Report, Pau, Rat-Stats, Venezia on May 7, 2019 by xi'an
From the recently started ASA books series on statistical reasoning in science and society (of which I already reviewed a sequel to The Lady tasting Tea), a short book, Statistics and Health Care Fraud, I read at the doctor while waiting for my appointment, with no chances of cheating! While making me realise that there is a significant amount of health care fraud in the US, of which I had never though of before (!), with possibly specific statistical features to the problem, besides the use of extreme value theory, I did not find me insight there on the techniques used to detect these frauds, besides the accumulation of Florida and Texas examples. As such this is a very light introduction to the topic, whose intended audience of choice remains unclear to me. It is stopping short of making a case for statistics and modelling against more machine-learning options. And does not seem to mention false positives… That is, the inevitable occurrence of some doctors or hospitals being above the median costs! (A point I remember David Spiegelhalter making a long while ago, during a memorable French statistical meeting in Pau.) The book also illustrates the use of a free auditing software called Rat-stats for multistage sampling, which apparently does not go beyond selecting claims at random according to their amount. Without learning from past data. (I also wonder if the criminals can reduce the chances of being caught by using this software.)
A second book on the “same” topic!, Measuring Crime, I read, not waiting at the police station, but while flying to Venezia. As indicated by the title, this is about measuring crime, with a lot of emphasis on surveys and census and the potential measurement errors at different levels of surveying or censusing… Again very little on statistical methodology, apart from questioning the data, the mode of surveying, crossing different sources, and establishing the impact of the way questions are stated, but also little on bias and the impact of policing and preventing AIs, as discussed in Weapons of Math Destruction and in some of Kristin Lum’s papers.Except for the almost obligatory reference to Minority Report. The book also concludes on an history chapter centred at Edith Abbott setting the bases for serious crime data collection in the 1920’s.
[And the usual disclaimer applies, namely that this bicephalic review is likely to appear later in CHANCE, in my book reviews column.]
