Archive for the R Category

the polls weren’t wrong [book review]

Posted in Books, R, Statistics, University life with tags , , , , , , , , , , , , , , , , on November 1, 2024 by xi'an

While Nate Silver and his colleagues (as The New York Times Nate Cohn) have brought (part of) the general public to adopt a (more) scientific perspective on political polls, the way to combine them, and the need to keep uncertainty fully quantified (witness this recent “Two Theories for Why the Polls Failed in 2020, and What It Means for 2024” by Nate Cohn for the NYT’s Tilt), the author of this book embarks upon a crusade (and a lengthy rant) against pollsters and analysts and media reporters, with the single, many times repeated, argument that non-responses and undecided voters are crucial for the final election outcome… And that a poll gives a snapshot of the current (time) population opinion, not a prediction of its future state. There is not the slightest trace of statistical depth, there is actually no statistics at all found throughout the book, apart from a section (p.281) entitled “Threats to inferential statistics” (with not no maths either, as a few ratio manipulations and a chapter title invoking Jakob Bernoulli!) do not count as maths!, but a lot of repetitions on the same theme and dismissal of statisticians’ analyses, like Nate Silver’s. Opposing to them the theories of Nick Panagakis, a 1990’s pollster. (Funny enough, the Amazon reviews include one “expert in inferential statistics, the major tool employed by Carl [Alen] in this book” and another one stating that “Carl Allen takes the reader through a journey towards statistical literacy“!) And an R code (p.52) for plotting the outcome of 30 Binomial random draws.

“Unless my efforts achieve far more notoriety than even my most optimistic forecast would predict, [the Proportional Method] is unlikely to go away any time soon.” p.183

“If it seems like I’m picking on FiveThirtyEight a lot, it’s not because there are no other forecasters who are better or worse.” p.235

“Sounding eerily like myself, [Nate Silver] pointed out [in 2008] that `many things can happen’ months before the election.” p.208

Reading through the book (during a trip to & from Warwick) was painful, both for the feeling of being stuck in a plane with a perfect unknown, next seat, trying to force their weird theory upon you and no way to escape their rant, as well as for the terrible style of said book, full of repetitions and one-sentence paragraphs. (Of course, nothing as bad as this time near the 2012 US elections I flew to Des Moines next to an inebriated woman that would not stop blathering about her life!) Or as I imagine a card game addict defending their martingale as a sure way to win against the casino. It is also the first time I see references repeated (in postcripts) as many times as they are cited within a chapter. There is no true insight on how polling companies construct their polling samples, how they post-process outcomes by regression techniques, and no reflection on the unique weirdness of the US electoral system in that a few States determine the outcome (rather than majority votes) and thus how a tiny number of voters (escaping the law of Large Numbers) hold the overall result in their hand.

Thus (as most readers will have forecasted) concluding by not recommending the book!

[Disclaimer about potential self-plagiarism: this post or an edited version of it could possibly appear in my Books Review section in CHANCE. Most unlikely though!]

Ali–Mikhail–Haq copula, [re]simulated

Posted in Books, R, Statistics with tags , , , , , , , , , on October 13, 2024 by xi'an

When looking for a copula I could simulate from (rather than the Gaussian copula), I found an algorithm for the Ali–Mikhail–Haq copula

C_\theta(u,v) = \frac{uv}{1-\theta(1-u)(1-v)}\quad -1<\theta<1

that was proposed by Kumar (2010) as reproduced above. But the method seriously fails in that the range of (U,V) resulting from the simulation does not even cover the (0,1)² square! Unless I made an R coding mistake (which is always a possibility).

sim=function(T=1e3,h=.5){ 
  o=matrix(runif(2*T),T,2)
  a=1-o[,1];b=1-h*(1+2*a*o[,2])+2*h^2*a^2*o[,2]
  d=1+h*(2-4*a+4*a*o[,2])+h^2*(1-4*a*o[,2] +4*a^2*o[,2])
  o[,2]=1-2*o[,2]*(a*h-1)^2/(b+sqrt(d))
  return(o)}

There is no explanation in the paper as to why this algorithm is (not!) working, besides the inverse cdf argument—with the reference in R copBasic temporarily worrying until I checked the cdf inversion is completely numerical—, but a correct version can be derived from inverting the conditional cdf of one component, V, given the other, U. Namely, since the conditional cdf is given by

F(v|U=u) = \frac{v(1-\theta(1-v))}{(1-\theta(1-u)(1-v))^2}

which leads to a second degree polynomial equation (in v) when solving the equation F(v|U=u) = w.

sim=function(T=1e3,h=.5){
  o=matrix(runif(2*T),T,2)
  v1=o[,1];w=o[,2]
  d=(2*h*v1*w-1-h)^2-4*(1-w)*h*(1-h*w*v1^2)
  o[,2]=-2*h*v1*w+1+h-sqrt(d))/(2*h*(1-h*w*v1^2)
  return(1-o)}

And with a more likely outcome (Xed checked by comparing F(u,v) with its empirical version for several pairs (u,v)):

Hands-On Differential Privacy [book review]

Posted in Books, R, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , on October 2, 2024 by xi'an

Hands-On Differential Privacy was published just a few months ago (from September  2024!) by (the US publisher) O’Reilly, famous for its programming and technical books with animal covers! A slate pencil sea urchin in the present case. The book is indeed classical O’Reilly’s, with lots of notes, little theory (or maths!) and symbols, a loose structuring of the chapters (no section numbers) and highly detailed examples, and of course plenty of OpenDP code inserts. For instance, in the present case, a case study about the privatization of a sample average x̄ that takes about ten pages. Terrible equation rendering btw (what’s wrong with LATEX?!).  Overall, I am quickly lost in most of the chapters due to a lack of a driving narrative, facing instead a catalogue of possible scenari and procedures, appearing one after the other as in a fashion show.

Hands-On Differential Privacy is written by Ethan Cowan, Michael Shoemate, and Mayana Pereira. I came across the book during the OpenDP workshop at Harvard [that took place right after my return from the Pacific Northwest] and it is definitely linked with OpenDP, all authors being  actually involved at one stage or another in the OpenDP Team. The style of the book is once again in tune with the O’Reilly manuals, which sort of clashes with my preferences. For instance, the introduction of differential privacy (Chapter 2) is quite extensive. Chapter 3 proceeds to teach about private data transform(ation)s, stability (a rewording of Lipschitz-ianity), with code illustrations, often repeating the earlier derivation (see eg p203), while Chapter 4 is its equivalent for private mechanisms. (With the diagrams Figures 3-1 and 4-1 differing only in highlighting/bolding different functions in a privatized data processing pipeline.) Returning to differential privacy with a privacy loss parameter and to Laplace and exponential mechanisms, Chapter 5 proposes several notions of privacy, all closed under post-processing. This includes Wasserman and Zhou (2010) interpretation of privacy as hypothesis testing, except it is not exploited further than connecting type I and type II with (ε,δ) parameters. Chapter 6 concludes Part I about concepts with a series of (fearless) combinators, keeping stability and privacy. With an increasing proportion of coding excerpts which I [imho] did not find particularly helpful.

Nothing about statistical loss of information or efficiency, bias, &tc. until Chapter 8 (p199) and even then so little. Part II is about practice, with a first Chapter  7 on setting a privacy unit (e.g., a person-month) before ensuring their privacy is protected. And discussing unbounded contributions (not unbounded data!). While Chapter 8 very thinly covers statistical modelling, while remaining agnostic about the choice of statistical procedures (Bayes being solely and naïvely mentioned for classification, furthermore with data-based evaluation of the class “prior” probabilities, p211). At this stage, procedures are often only defined through spinets of code, like the private Theil-Sen estimator (pp204-205). The continuous case boils to a Normality assumption, with its pmf being defined (p212) as

\text{Pr}(x=\mu)=\frac{1}{\sqrt{2\pi\sigma}}e^{-(x-\mu)/2\sigma^2}

which contains at least three errors! Chapter 9 is the equivalent of Chapter 8 for machine learning, mostly centred on private gradient descent. And a Pytorch section (pp232-235). Completed by a light Chapter 10 on synthetic data, which does not seem to broach upon the issue of large dimension covariates, providing instead a list of GAN synthetizers.

Part III (Deploying differential privacy) is even more about practice, with Chapter 11 on privacy attacks, Chapter 12 on calibrating a privacy mechanism (co-written with Jayshree Sarathy), and good practice (like codebooks and data annotations), with the appearance of contextual integrity I discovered if not perfectly understood last year at the BIRS workshop in Kelowna. And Chapter 13 on planning a privacy project, with an 11 step checklist, most of which are quite vague [imho] and do include strategies to make the data owners confident their privacy is safe.

[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE]

joint fiddlin

Posted in Books, Kids, R, Statistics with tags , , , , , on April 22, 2024 by xi'an

Flip a fair coin 100 times, resulting in a sequence of heads (H) and tails (T). For each HH in the sequence, Alice gets a point; for each HT, Bob does, so e.g. for the subsequence THHHT Alice gets 2 points and Bob gets 1 point. Who is most likely to win?

An interesting conundrum in that the joint distribution of (A,B) need be considered for showing that Bob is more likely. Indeed, looking at the marginals does not help since the probability of the base events is the same. A solution on X validated (for a question posted when the Fiddler’s puzzle came out, Friday morn) demonstrates via a four state Markov chain representation the result (obvious from a quick simulation) that Alice wins 45% of the time while Bob wins 48%. The intuition is that, each time Alice wins at least a point, Bob gets an extra point at the end of the sequence (except possibly at the stopping time t=100), while in other cases Alice and Bob have the same probability to win one point.

python [book review]

Posted in Books, Kids, R with tags , , , , , , , , , , , on April 19, 2024 by xi'an

A fellow coder shared with me this recent manual (in French) entitled python (for the computer language, not the snake) written by Nathalie Azoulai  as he found it an interesting literary (if not computer) program. It parses rather quickly and I compiled it in one single run on my way to Bristol [Mecca of punched card coders!] last week. The core idea of this manual is one of a middle-aged, clueless (about coding), single-mother, writer engaging into an experiment to understand coding and coders. She does not succeed. And the manual also fails to implement said idea, at least for me. This writer calls upon a series of tutors to get into understanding coding and learning python but she does not get much further than exiting the “hello world” instruction that seem to bug the early chapters of most programming manuals! It sounds as if the compiler (character or author) is mostly uninterested in the fundamental function of coding towards delivering, that is, that it is used and useful for solving real problems. Instead, the manual leans towards a massive and falsely naïve infodump on the history of coding, with very few surprises there, along with a gradually overwhelming, personal, and frankly of very limited interest, introspection on a past and ambiguous love (or sexual attraction) that returned too many error messages. The end of the manual is a predictable mise-en-abîme where the writer finds Chat-GPT unsurprisingly lacking in its ability to deliver a novel that goes beyond the (re)production of earlier codes, e.g., clichés and conventional situations. I would certainly not recommend this shallow, unrealistic, and inherently flawed, python manual to anyone…

Since I mentioned ChatGPT, here is its output when repeatedly insisting on a computer manual style, but funnily missing some obvious programming synonyms:

Received from a fellow enthusiast, this recent manuscript by Nathalie Azoulai is touted as a captivating computational experiment. It swiftly parses, and I sequentially traversed its contents during my recent journey, processing it in one uninterrupted session. At its core, the narrative embarks on an intriguing quest to decrypt the intricacies of programming paradigms and the programmer’s mindset. However, its execution fails to trigger the desired engagement, encountering crashes and bugs along the way. Azoulai’s documentation chronicles her [character’s] endeavor to master Python, facilitated by a chain of mentors who, regrettably, encounter runtime errors in their attempts to elevate her beyond the rudimentary “hello world” commands—a fundamental subroutine in the lexicon of programming vernacular. The narrative raises questions about the intrinsic motivations driving one’s pursuit of programming prowess, encountering recursive loops of introspection. Rather than delving into the depths of coding intricacies, the manuscript deviates into an extensive discourse on the historical evolution of programming—a divergence that, while intellectually stimulating, detracts from the central narrative focus. Interspersed amidst this cerebral exploration are fragments of introspection, reflections of past affections interwoven delicately into the fabric of Azoulai’s discourse, yet they fail to meet the conditionals of relevance. Yet, as the narrative unfolds, it descends into a labyrinth of self-reflection, where the author grapples with the constraints of artificial intelligence—a poignant reflection on the interplay between creativity and the digital frontier, encountering exception handling challenges along the way. In conclusion, Azoulai’s manuscript presents an intriguing blend of literary and computational exploration, yet falls short of achieving transcendence. While it beckons towards the precipice of profound inquiry, its meandering narrative trajectory ultimately leaves the reader yearning for a more cohesive and resonant exploration of its thematic landscape, perhaps requiring debugging and refactoring for optimal performance.

ChatGPT also produced a python code I do not want to inflict in its entirety upon ‘Og’s readers…