Archive for accountability

Data science ethics [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on May 5, 2025 by xi'an

Data science ethics (concepts, techniques and cautionary tales), by David Martens, was published in 2022 by Oxford University Press. The book is inspired by the author’s  course on Data Science and ethics he has been teaching at the University of Antwerp. (With a link to his slides.) The 255p book proceeds by decomposing the ethics of data science into its different steps: data gathering (Chap. 2), data preprocessing (Chap. 3), modelling (Chap. 4), evaluation (Chap. 5), and deployment (Chap. 6). Following the `FAT Flow Framework´, where FAT stands for fairness, accountability, and transparency.

Do not expect much maths, stats, or anything quantitative: this book is mostly about concepts, even though some (mostly well-known) illustrations are provided. Chapter 2 includes a description of encryption (with homomorphic encryption treated in Chapter 4). And somewhat improbably quantum computing. Differential privacy gets a few pages with not a single formula (until Chapter 4, again).

Chapter 3 covers k-anonymity, record linkage, reidentification, (through cautionary tales) and discrimination through biases in the (learning) dataset.  Chapter 4 is defining ε differential privacy with the Laplace randomization as a possible implementation and with no critical stance on the limitations of the concept. The computation limitations of homomorphic encryption are more clearly pointed out. Federated learning is only quickly mentioned. The section about measuring fairness and reducing bias implies that some prior knowledge is available about whom is potentially discriminated and which covariates to add to the model. The last section on explicability of predictions is worthwhile in signalling the difficulty with most (black box) AI but the example opposing an SVM model to a logistic model is not tremendously convincing in that neither model is true.

Chapter 5 addresses the crucial challenge of ethical evaluation in a rather verbose and vague manner. For instance, with no instruction on how to resist adversarial attacks. Or criticising p-hacking and multiple testing while missing the elephant in the room (p-values!). Drifting from the topic when discussing the misdeeds of Diederik Stapel. Most of the same goes about Chapter 6 and its take on ethical deployment, when going through examples such as Google’s policies in China. Or general musing on the impact of AI on societal inequalities (with mentions of companies and CEOs who have since then back-pedalled on their ethical engagement). These chapters are lacking in tools and (more) practical recommendations.

One interesting aspect of the book is the attention paid to the EU(ropean) aspect of these concerns, through the GDPR (General DAta Protection Regulations) adopted by the European Parliament in 2016. (There is also a brief mention of China’s regulations, but no details beyond a reference. Maybe the Chinese edition differs.)

the privacy fallacy [book review]

Posted in Statistics with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on May 3, 2024 by xi'an

“The World changed significantly since 1973.” (p.10)

I read this book, The Privacy Fallacy: Harm and Power in the Information Economy, by Ignacio Cofone, upon my return from Warwick the past week. This is a Cambridge University Press 2023 book I had picked from their publication list after reviewing a book proposal for them. A selection made with our ERC OCEAN goals in mind, but without paying enough attention to the book table of contents, since it proved to be a Law book!

“People’s inability to assess privacy risks impact people’s behavior toward privacy because it turns the risks into uncertainty, a kind of risk that is impossible to estimate.” (p.31)

Still, this ended up being a fairly interesting read (for me) about the shortcomings of the current legal privacy laws (in various countries), since they are based on an obsolete perception that predates AIs and social media. Its main theme is that privacy is a social value that must be protected, regardless of whether or not its breach has tangible consequences. The author then argues that notions that support these laws such as the rationality of individual choices, the confusion between privacy and secrecy, the binary dichotomy between public and private, &tc., all are erroneous, hence the “fallacy” he denounces. One immediate argument for his position is the extreme imbalance of information between individuals and corporations, the former being unable to assess the whole impact of clicking on “I agree” when visiting a webpage or installing a new app. The more because the data thus gathered is pipelined to third parties. (“One’s efforts cannot scale to the number of corporations collecting and using one’s personal data”, p.93) For similar reasons, Cofone further states that the current principles based on contracts are inappropriate. Also because data harm can be collective and because companies have a strong incentive to data exploitation, hence a moral hazard.

“Inferences, relational data, and de-identified data aren’t captured by consent provisions.” (p.9)

“AI inferences worsen information overload (…) As [they] continue to grow, so will the insufficiency of our processing ability to estimate our losses.” (p.75)

As illustrated by the surrounding quotes, the statistical and machine-learning aspects of the book are few and vague, in that the additional level of privacy loss due to post-data processing is considered as a further argument for said loss to be impossible to quantify and assess, without a proper evaluation of the channels through which this can happen and without a reglementary proposal towards its control. This level of discourse makes AIs appear as omniscient methods, unfortunately.

“Inferences are invisible (…) Risks posed by inferences are impossible to anticipate because the information inferred is disproportionate to the sum of the information disclosed.” (p.49)

“The idea of probabilistic privacy loss is crucial in a world where entities (..) mostly affect our privacy by making inferences” (p.121)

The attempts at regulation such as opt-in and informed consent are then denounced as illusions—obviously so imho, even without considering the nuisance of having to click on “Reject” for each newly visited website!—. De- and re-identified data does not require anyone’s consent. Data protection rights, as of today, do not provide protection in most cases, the burden of proof residing on the privacy victims rather than the perpetrators. The book unsurprisingly offers no technical suggestion towards ensuring corporations and data brokers comply with this respect of privacy and on the opposite agrees that institutional attempts such as GDPR remain well-intended wishful thinking w/o imposing a hard-wired way of controlling the data flows, with the “need of an enforcement authority with investigating and sanctioning powers” (p.106) . The only in-depth proposal therein is pushing for stronger accountability of these corporations via a new type of liability, with a prospect of class actions (if only in countries with this judiciary possibility).

[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE.]

algorithm for predicting when kids are in danger [guest post]

Posted in Books, Kids, Statistics with tags , , , , , , , , , , , , , , , , , on January 23, 2018 by xi'an

[Last week, I read this article in The New York Times about child abuse prediction software and approached Kristian Lum, of HRDAG, for her opinion on the approach, possibly for a guest post which she kindly and quickly provided!]

A week or so ago, an article about the use of statistical models to predict child abuse was published in the New York Times. The article recounts a heart-breaking story of two young boys who died in a fire due to parental neglect. Despite the fact that social services had received “numerous calls” to report the family, human screeners had not regarded the reports as meeting the criteria to warrant a full investigation. Offered as a solution to imperfect and potentially biased human screeners is the use of computer models that compile data from a variety of sources (jails, alcohol and drug treatment centers, etc.) to output a predicted risk score. The implication here is that had the human screeners had access to such technology, the software might issued a warning that the case was high risk and, based on this warning, the screener might have sent out investigators to intervene, thus saving the children.

These types of models bring up all sorts of interesting questions regarding fairness, equity, transparency, and accountability (which, by the way, are an exciting area of statistical research that I hope some readers here will take up!). For example, most risk assessment models that I have seen are just logistic regressions of [characteristics] on [indicator of undesirable outcome]. In this case, the outcome is likely an indicator of whether child abuse had been determined to take place in the home or not. This raises the issue of whether past determinations of abuse– which make up  the training data that is used to make the risk assessment tool–  are objective, or whether they encode systemic bias against certain groups that will be passed through the tool to result in systematically biased predictions. To quote the article, “All of the data on which the algorithm is based is biased. Black children are, relatively speaking, over-surveilled in our systems, and white children are under-surveilled.” And one need not look further than the same news outlet to find cases in which there have been egregiously unfair determinations of abuse, which disproportionately impact poor and minority communities.  Child abuse isn’t my immediate area of expertise, and so I can’t responsibly comment on whether these types of cases are prevalent enough that the bias they introduce will swamp the utility of the tool.

At the end of the day, we obviously want to prevent all instances of child abuse, and this tool seems to get a lot of things right in terms of transparency and responsible use. And according to the original article, it (at least on the surface) seems to be effective at more efficiently allocating scarce resources to investigate reports of child abuse. As these types of models become used more and more for a wider variety of prediction types, we need to be cognizant that (to quote my brilliant colleague, Josh Norkin) we don’t “lose sight of the fact that because this system is so broken all we are doing is finding new ways to sort our country’s poorest citizens. What we should be finding are new ways to lift people out of poverty.”