Archive for slides

Data science ethics [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on May 5, 2025 by xi'an

Data science ethics (concepts, techniques and cautionary tales), by David Martens, was published in 2022 by Oxford University Press. The book is inspired by the author’s  course on Data Science and ethics he has been teaching at the University of Antwerp. (With a link to his slides.) The 255p book proceeds by decomposing the ethics of data science into its different steps: data gathering (Chap. 2), data preprocessing (Chap. 3), modelling (Chap. 4), evaluation (Chap. 5), and deployment (Chap. 6). Following the `FAT Flow Framework´, where FAT stands for fairness, accountability, and transparency.

Do not expect much maths, stats, or anything quantitative: this book is mostly about concepts, even though some (mostly well-known) illustrations are provided. Chapter 2 includes a description of encryption (with homomorphic encryption treated in Chapter 4). And somewhat improbably quantum computing. Differential privacy gets a few pages with not a single formula (until Chapter 4, again).

Chapter 3 covers k-anonymity, record linkage, reidentification, (through cautionary tales) and discrimination through biases in the (learning) dataset.  Chapter 4 is defining ε differential privacy with the Laplace randomization as a possible implementation and with no critical stance on the limitations of the concept. The computation limitations of homomorphic encryption are more clearly pointed out. Federated learning is only quickly mentioned. The section about measuring fairness and reducing bias implies that some prior knowledge is available about whom is potentially discriminated and which covariates to add to the model. The last section on explicability of predictions is worthwhile in signalling the difficulty with most (black box) AI but the example opposing an SVM model to a logistic model is not tremendously convincing in that neither model is true.

Chapter 5 addresses the crucial challenge of ethical evaluation in a rather verbose and vague manner. For instance, with no instruction on how to resist adversarial attacks. Or criticising p-hacking and multiple testing while missing the elephant in the room (p-values!). Drifting from the topic when discussing the misdeeds of Diederik Stapel. Most of the same goes about Chapter 6 and its take on ethical deployment, when going through examples such as Google’s policies in China. Or general musing on the impact of AI on societal inequalities (with mentions of companies and CEOs who have since then back-pedalled on their ethical engagement). These chapters are lacking in tools and (more) practical recommendations.

One interesting aspect of the book is the attention paid to the EU(ropean) aspect of these concerns, through the GDPR (General DAta Protection Regulations) adopted by the European Parliament in 2016. (There is also a brief mention of China’s regulations, but no details beyond a reference. Maybe the Chinese edition differs.)

a minor typo, from 1986

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , on December 2, 2024 by xi'an

venISBA⁴⁻

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , on July 8, 2024 by xi'an

As I was released all of a sudden from the Ospedale Civile di Venezia around noon, I managed to attend the last session of ISBA 2024 (after stopping by my airbnb for an emergency coffee next to the hospital and stopping for showering, changing clothes, and eating something more substantial than the contents of IV bags).

My first of these last talks was about coresets by Trevor Campbell, for reducing sample sizes while keeping the likelihood roughly the same (and making me wondering if possibly getting some privacy on the side??) Original algorithm almost completely blind to the data, but a new version by subsample-optimize (KL distance to the posterior) version bringing huge improvements (although I missed the practical details on how the algorithm is reaching this minimum), namely a KL distance of order O(1), i.e., not growing in the sample size. Then, in the same session, a talk by Aikihiko Nakamura on mixing and PDMP, resulting in the novel bouncy Hamiltonian dynamics, which proves time reversible and volume preserving, with no U turns and the time within a given general Hamiltonian value being itself generated w/o rejection. (I am quite sorry to have missed other PDMP talks during the conference, eg, Paul Fearnhead’s, as well as the last poster session…) And I finally jumped rooms to listen to Sam Power on hybrid slice sampling with an MCMC extension to avoid simulating from the Uniform conditional. Reminding me of nested sampling, which also faces this difficulty of sampling from a possibly complex set. This was the end of a wonderful (if shortened by my personal issue) meeting. Next round, see you in Nagoya, Japan (on the Tōkaidō road!).


As a final word about this ISBA 2024 conference in Ca’Foscari, on many levels, I want to most warmly thank my friend Roberto Casarin for his investment and dedication for making the event running so efficiently, in an ideal environment for a meeting of this (800+) size that kept to the Aristotelian unities, especially keeping people together on a unique site without feeling crowded (and very few falling in a Venice canal). And many thanks as well to the local organisers (discounting my nominal inclusion in that group!), the Ca’Foscari staff, and all the students involved in the event!

repulsive sampling

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , on January 31, 2024 by xi'an

After a long absence from the monthly Séminaire Parisien de Statistique I attended one today at IHP, including a talk by Diala Hawat on repelled point processes for numerical integration by Hawat et al. The goal is to get (and prove) a universal variance improvement for numerical integration by applying a form of determinantal processes to initial simulations, as eg iid (Poisson process) sampling (without accounting for the O(N²) cost in moving these points). The repelled points are obtain by a single (why single?) move based on a force function (as shown in the slide below), inspired by a Coulomb potential (in the sense that said move appears as one gradient step along the potential). Which reminded me of the pinball sampler, even though the inverse norm was just there to create infinite repulsion near each point. A surprising feature of this repelling step is that it even modifies a (QMC) Sobol process with also an (empirical) improvement in the variance. I wonder if one could construct an MCMC algorithm that would target a joint distribution, maybe via a copula representation, maybe via an equivalent version of HMC.


As an aside, the Bakhvalov results on the existence of a worst case integrand for any deterministic or random sequence (see top slide) made me wonder what the shape of this worst case function is, esp. for a QMC sequence (eg, Sobol). And whether or not they are of any relevance as a counterfactor to the optimal importance functions.

mixtures at BNP [slides]

Posted in Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , on October 29, 2022 by xi'an

After chatting with some BNP13 participants at the Puerto Montt airport, I gave in (!) to their kind request to put my slides on-line and here is the link to the slideshare depository. It was quite the nice coincidence that Sanjib Basu (whom I met in Purdue in 1987!) gave the invited talk in our session since we were building on the under-appreciated Basu-Chib approximation of the evidence. Overall, this was an exhilarating week and I now have to recover from this sensory overload. (Incidentally, and uninterestingly, I got swindled by not one but two taxis on my way back to Santiago!)