This semester, I—as a teacher—came across two cases of heavily reliance on AI by master students, mostly for coding purposes, to which I had rather surprisingly not been exposed before. (Except for this plagiarised thesis two years ago that essentially rewrote existing papers with synonyms and for which we had to get to the disciplinary committee!) One project made a massive advance within two days, with hundreds of lines of beautiful python code, and reasonable output, but with my student unable to explain the code or the method behind… And anther case homeworks involving coding came back with extremely clean codes as well. Meaning they could not be graded and we had to switch to another type of evaluation. Oh well, welcome ol’me into the new age (just for a few years!)
Archive for Python
fAIrst contAIct
Posted in Books, Kids, pictures, Statistics, University life with tags ChatGPT, coding, debugging, discipline, master project, plagiarism, Python, R, teaching, Université Paris Dauphine, University of Warwick on November 12, 2025 by xi'ancomputo on the go
Posted in Books, R, Statistics, University life with tags binary, Biometrika, Computo, editor in chief, free software, French researchers, github, journal, Julia, Latin, logo, machine learning, open access, open source, Python, Quarto, R, repositories, reproducible research, Rmarkdown, SFDS, Société française de Statistique, Statistics on April 8, 2025 by xi'an
Hands-On Differential Privacy [book review]
Posted in Books, R, Statistics, University life with tags Bayesian inference, biases, book review, CHANCE, classification, coding, contextual integrity, differential privacy, GANs, Harvard, Laplace distribution, LaTeX, Lipschitz continuity, machine learning, O'Reilly Media, OpenDP, Python, slate pencil sea urchin, Type I error, Type II error on October 2, 2024 by xi'anHands-On Differential Privacy was published just a few months ago (from September 2024!) by (the US publisher) O’Reilly, famous for its programming and technical books with animal covers! A slate pencil sea urchin in the present case. The book is indeed classical O’Reilly’s, with lots of notes, little theory (or maths!) and symbols, a loose structuring of the chapters (no section numbers) and highly detailed examples, and of course plenty of OpenDP code inserts. For instance, in the present case, a case study about the privatization of a sample average x̄ that takes about ten pages. Terrible equation rendering btw (what’s wrong with LATEX?!). Overall, I am quickly lost in most of the chapters due to a lack of a driving narrative, facing instead a catalogue of possible scenari and procedures, appearing one after the other as in a fashion show.
Hands-On Differential Privacy is written by Ethan Cowan, Michael Shoemate, and Mayana Pereira. I came across the book during the OpenDP workshop at Harvard [that took place right after my return from the Pacific Northwest] and it is definitely linked with OpenDP, all authors being actually involved at one stage or another in the OpenDP Team. The style of the book is once again in tune with the O’Reilly manuals, which sort of clashes with my preferences. For instance, the introduction of differential privacy (Chapter 2) is quite extensive. Chapter 3 proceeds to teach about private data transform(ation)s, stability (a rewording of Lipschitz-ianity), with code illustrations, often repeating the earlier derivation (see eg p203), while Chapter 4 is its equivalent for private mechanisms. (With the diagrams Figures 3-1 and 4-1 differing only in highlighting/bolding different functions in a privatized data processing pipeline.) Returning to differential privacy with a privacy loss parameter and to Laplace and exponential mechanisms, Chapter 5 proposes several notions of privacy, all closed under post-processing. This includes Wasserman and Zhou (2010) interpretation of privacy as hypothesis testing, except it is not exploited further than connecting type I and type II with (ε,δ) parameters. Chapter 6 concludes Part I about concepts with a series of (fearless) combinators, keeping stability and privacy. With an increasing proportion of coding excerpts which I [imho] did not find particularly helpful.
Nothing about statistical loss of information or efficiency, bias, &tc. until Chapter 8 (p199) and even then so little. Part II is about practice, with a first Chapter 7 on setting a privacy unit (e.g., a person-month) before ensuring their privacy is protected. And discussing unbounded contributions (not unbounded data!). While Chapter 8 very thinly covers statistical modelling, while remaining agnostic about the choice of statistical procedures (Bayes being solely and naïvely mentioned for classification, furthermore with data-based evaluation of the class “prior” probabilities, p211). At this stage, procedures are often only defined through spinets of code, like the private Theil-Sen estimator (pp204-205). The continuous case boils to a Normality assumption, with its pmf being defined (p212) as
which contains at least three errors! Chapter 9 is the equivalent of Chapter 8 for machine learning, mostly centred on private gradient descent. And a Pytorch section (pp232-235). Completed by a light Chapter 10 on synthetic data, which does not seem to broach upon the issue of large dimension covariates, providing instead a list of GAN synthetizers.
Part III (Deploying differential privacy) is even more about practice, with Chapter 11 on privacy attacks, Chapter 12 on calibrating a privacy mechanism (co-written with Jayshree Sarathy), and good practice (like codebooks and data annotations), with the appearance of contextual integrity I discovered if not perfectly understood last year at the BIRS workshop in Kelowna. And Chapter 13 on planning a privacy project, with an 11 step checklist, most of which are quite vague [imho] and do include strategies to make the data owners confident their privacy is safe.
[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE]
python [book review]
Posted in Books, Kids, R with tags book review, Bristol, Bristol cardboard, ChatGPT, coding, computer language, computer science, French literature, P.O.L., programming style, punched card, Python on April 19, 2024 by xi'an
A fellow coder shared with me this recent manual (in French) entitled python (for the computer language, not the snake) written by Nathalie Azoulai as he found it an interesting literary (if not computer) program. It parses rather quickly and I compiled it in one single run on my way to Bristol [Mecca of punched card coders!] last week. The core idea of this manual is one of a middle-aged, clueless (about coding), single-mother, writer engaging into an experiment to understand coding and coders. She does not succeed. And the manual also fails to implement said idea, at least for me. This writer calls upon a series of tutors to get into understanding coding and learning python but she does not get much further than exiting the “hello world” instruction that seem to bug the early chapters of most programming manuals! It sounds as if the compiler (character or author) is mostly uninterested in the fundamental function of coding towards delivering, that is, that it is used and useful for solving real problems. Instead, the manual leans towards a massive and falsely naïve infodump on the history of coding, with very few surprises there, along with a gradually overwhelming, personal, and frankly of very limited interest, introspection on a past and ambiguous love (or sexual attraction) that returned too many error messages. The end of the manual is a predictable mise-en-abîme where the writer finds Chat-GPT unsurprisingly lacking in its ability to deliver a novel that goes beyond the (re)production of earlier codes, e.g., clichés and conventional situations. I would certainly not recommend this shallow, unrealistic, and inherently flawed, python manual to anyone…
Since I mentioned ChatGPT, here is its output when repeatedly insisting on a computer manual style, but funnily missing some obvious programming synonyms:
Received from a fellow enthusiast, this recent manuscript by Nathalie Azoulai is touted as a captivating computational experiment. It swiftly parses, and I sequentially traversed its contents during my recent journey, processing it in one uninterrupted session. At its core, the narrative embarks on an intriguing quest to decrypt the intricacies of programming paradigms and the programmer’s mindset. However, its execution fails to trigger the desired engagement, encountering crashes and bugs along the way. Azoulai’s documentation chronicles her [character’s] endeavor to master Python, facilitated by a chain of mentors who, regrettably, encounter runtime errors in their attempts to elevate her beyond the rudimentary “hello world” commands—a fundamental subroutine in the lexicon of programming vernacular. The narrative raises questions about the intrinsic motivations driving one’s pursuit of programming prowess, encountering recursive loops of introspection. Rather than delving into the depths of coding intricacies, the manuscript deviates into an extensive discourse on the historical evolution of programming—a divergence that, while intellectually stimulating, detracts from the central narrative focus. Interspersed amidst this cerebral exploration are fragments of introspection, reflections of past affections interwoven delicately into the fabric of Azoulai’s discourse, yet they fail to meet the conditionals of relevance. Yet, as the narrative unfolds, it descends into a labyrinth of self-reflection, where the author grapples with the constraints of artificial intelligence—a poignant reflection on the interplay between creativity and the digital frontier, encountering exception handling challenges along the way. In conclusion, Azoulai’s manuscript presents an intriguing blend of literary and computational exploration, yet falls short of achieving transcendence. While it beckons towards the precipice of profound inquiry, its meandering narrative trajectory ultimately leaves the reader yearning for a more cohesive and resonant exploration of its thematic landscape, perhaps requiring debugging and refactoring for optimal performance.
ChatGPT also produced a python code I do not want to inflict in its entirety upon ‘Og’s readers…
combining normalizing flows and QMC
Posted in Books, Kids, Statistics with tags Arianna Rosenbluth, arXiv, Illya Sobol, importance sampling, inverse cdf, John Halton, MCM 2023, Metropolis-Hastings algorithm, mostly Monte Carlo seminar, normalizing flow, Python, quasi-Monte Carlo methods, scrambling, Sobol sequences on January 23, 2024 by xi'an
My PhD student Charly Andral [presented at the mostly Monte Carlo seminar and] arXived a new preprint yesterday, on training a normalizing flow network as an importance sampler (as in Gabrié et al.) or an independent Metropolis proposal, and exploiting its invertibility to call quasi-Monte Carlo low discrepancy sequences to boost its efficiency. (Training the flow is not covered by the paper.) This extends the recent study of He et al. (which was presented at MCM 2023 in Paris) to the normalising flow setting. In the current experiments, the randomized QMC samples are computed using the SciPy package (Roy et al. 2023), where the Sobol’ sequence is based on Joe and Kuo (2008) and on Matouˇsek (1998) for the scrambling, and where the Halton sequence is based on Owen (2017). (No pure QMC was harmed in the process!) The flows are constructed using the package FlowMC. As expected the QMC version brings a significant improvement in the quality of the Monte Carlo approximations, for equivalent computing times, with however a rapid decrease in the efficiency as the dimension of the targetted distribution increases. On the other hand, the architecture of the flow demonstrates little relevance. And the type of RQMC sequence makes a difference, the advantage apparently going to a scrambled Sobol’ sequence.