Archive for contextual integrity

no privacy in nature

Posted in Books, Statistics, University life with tags , , , , , , , , , on September 26, 2026 by xi'an

A paper about privacy (or lack thereof) in Nature by Knolle et al. about medical AI models that exhibit a weakness to membership privacy attacks thru multiple queries of the public model. As discussed in a commentary article by Zhang & Ghassemi, a membership privacy attack proves successful when the confidence of the model prediction jumps to higher values for a real target of interest compared with an imaginary one. This obviously assume that users (and attackers) may repeat queries ad nauseam from the model, which differs from our Bayesian privacy setting where the output is provided once and only once (whatever the release mechanism is).

The attacker is modelled as resorting to a basic likelihood-ratio MIAs3,4, that is, a test based on the prediction confidence attached to the target model with the null being that the target is not a member:

“the parameters of the distributions under the two hypotheses are specified by parametric fitting of sample confidence values obtained from reference models. Reference models are models assumed to be trained by the attacker and are ideally, but not necessarily, of similar architecture as the target model and trained on data similar to the training dataset” – M Knolle & al.

but the paper does not provide further details about inferring about the model parameters. The following sentence is also unclear

“objectively larger threats are posed by privacy attacks with stronger assumptions on a potential attacker, such as access to model parameters17, access to parameter updates during model training18 (…) By contrast, the type of attack we consider here requires querying the target model only once (to obtain a prediction for the target record)” – M Knolle & al.

in that, indeed, returning the model parameter estimates is more informative about the data than a black-box prediction interface, but I miss the single query point as it seems to me that the attacker must multiply the queries to build a confidence distribution under both hypotheses.

“Purely technical measures, such as a mathematical approach called differential privacy9, often create performance trade-offs10 that are too limiting for medical AI tools. Instead, what is needed are regulatory and sociotechnical safeguards, such as privacy audits and risk assessments, that are specific to the domain in question.” – H Zhang & M Ghassemi

“our results indicate that privacy attacks against AI models may be much more effective at compromising the privacy of individual data contributors than previously thought. This suggests that current AI privacy risk reporting practices may underestimate individual-level risk and thus motivates the integration of mathematically verifiable risk mitigation strategies such as differential privacy (DP) into medical AI model development workflows.” – M Knolle & al.

“the finding that record-level differential privacy is insufficient for multi-record patients is particularly impactful and has clear policy and implementation implications” –Referee #2

“we also found that the full mitigation of MIAs for all data-contributing patients requires stricter levels of privacy protection (ϵ, δ ) smaller than previously believed. Moreover, our results also show that fully mitigating MIAs requires DP accounting at the patient- rather than the record-level.” – M Knolle & al.

In a funny (?) clash, the authors of the paper and those of the comment and a referee seem to disagree on the pertinence of differential privacy guarantees in this context. But the conclusion remains very vague on a rigorous way to assert privacy leaks and confidentiality protection for a given AI and its supporting dataset.

 

 

Bayesian Adversarial Privacy [v2]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , on September 11, 2026 by xi'an

We have just reposted our paper Bayesian Adversarial Privacy on arXiv to reflect the revision we wrote in the past months, to address the (quite sensible) comments from the reviewers. Interestingly the discussants of my Akaike lecture made similar points. The main changes are in explaining more clearly the nature of the combined loss, with Antoine coming up with the use of illuminating R-U map representations, in enlarging the references to other approaches, in stressing that Eve was an Alice’s construct rather than a genuine adversary, but still integrating the case of “multiple Eves”, in mellowing our criticisms of DP, and in expanding the conclusion with limitations and extensions subsections.

a journal of the chaos (en cuisine) year

Posted in Books, Kids, Mountains, pictures, Travel, Wines with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on June 12, 2025 by xi'an

 Read La Fille du Grand Hiver (The Daughter of the Great Winter) by Isabelle Autissier (also a sailor who was the first woman to complete a solo world race in 1991). This is a novelised version of the story of Arnarulunguaq, who accompanied Knud Rasmussen on the Fifth Thule Expedition over several years, all the way to the Bering Strait. She died of tuberculosis the year after. (And will appear on a Dane banknote soon.) The novel is told from Arnarulunguaq‘s perspective, as she opens to wider possibilities, cultural diversity, and traces back her Inuit lineage along Knud Rasmussen. It is a beautiful travel story (and then some more), in the fascinating Far North, which further reminded me of a catholic missionary visiting our primary school in the mid  1960’s and giving a conference on the Inuits and their culture. I pressed my parents until they bought his book and read it over and over, probably the seed of my fascination for the Far North!

Mis-made both a vat of rice pudding—presumably by first grinding the rice grains as the whey did not get absorbed—and a vat of chocolate mousse!—by not melting the chocolate long enough—and another of vegan mousse—by failing to notice the chickpea water was salted and that my MIL’s egg-beater was broken (but I salvaged it into a decent mole, if not with 49 ingredients)! I also grilled panned hot chilis I had found on a local market, but they were so hot they made me cough and open all windows; they are now resting in an olive oil jar, to be soon tested!!! Rhubarb is back in the stalls of my local market, providing me with my breakfast spread for the coming weeks. I again enjoyed Toukoul’s dora wat on the night of the privaCI workshopt—a workshop that focussed almost entirely on contextual integrity, if with very little overlap with the BIRS workshop in Kelowna. And harvested as many cherries as possible from our fruitful (!) cherry tree (#1) before the local birds are them all.

Watched Black Doves, with (Bend it Like Beckham) Keira Knightley as the only recognisable actress in the set, a very cartoony spies series taking place in London at Xmas time, with funny dialogues and too good vibes that should restrict a viewer to watching the show on 24 December and no other day.

The 7th Annual PrivaCI Symposium, 19-20 May, Brussels

Posted in pictures, Statistics, University life with tags , , , , , , , , , , , , , , on January 24, 2025 by xi'an

Hands-On Differential Privacy [book review]

Posted in Books, R, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , on October 2, 2024 by xi'an

Hands-On Differential Privacy was published just a few months ago (from September  2024!) by (the US publisher) O’Reilly, famous for its programming and technical books with animal covers! A slate pencil sea urchin in the present case. The book is indeed classical O’Reilly’s, with lots of notes, little theory (or maths!) and symbols, a loose structuring of the chapters (no section numbers) and highly detailed examples, and of course plenty of OpenDP code inserts. For instance, in the present case, a case study about the privatization of a sample average x̄ that takes about ten pages. Terrible equation rendering btw (what’s wrong with LATEX?!).  Overall, I am quickly lost in most of the chapters due to a lack of a driving narrative, facing instead a catalogue of possible scenari and procedures, appearing one after the other as in a fashion show.

Hands-On Differential Privacy is written by Ethan Cowan, Michael Shoemate, and Mayana Pereira. I came across the book during the OpenDP workshop at Harvard [that took place right after my return from the Pacific Northwest] and it is definitely linked with OpenDP, all authors being  actually involved at one stage or another in the OpenDP Team. The style of the book is once again in tune with the O’Reilly manuals, which sort of clashes with my preferences. For instance, the introduction of differential privacy (Chapter 2) is quite extensive. Chapter 3 proceeds to teach about private data transform(ation)s, stability (a rewording of Lipschitz-ianity), with code illustrations, often repeating the earlier derivation (see eg p203), while Chapter 4 is its equivalent for private mechanisms. (With the diagrams Figures 3-1 and 4-1 differing only in highlighting/bolding different functions in a privatized data processing pipeline.) Returning to differential privacy with a privacy loss parameter and to Laplace and exponential mechanisms, Chapter 5 proposes several notions of privacy, all closed under post-processing. This includes Wasserman and Zhou (2010) interpretation of privacy as hypothesis testing, except it is not exploited further than connecting type I and type II with (ε,δ) parameters. Chapter 6 concludes Part I about concepts with a series of (fearless) combinators, keeping stability and privacy. With an increasing proportion of coding excerpts which I [imho] did not find particularly helpful.

Nothing about statistical loss of information or efficiency, bias, &tc. until Chapter 8 (p199) and even then so little. Part II is about practice, with a first Chapter  7 on setting a privacy unit (e.g., a person-month) before ensuring their privacy is protected. And discussing unbounded contributions (not unbounded data!). While Chapter 8 very thinly covers statistical modelling, while remaining agnostic about the choice of statistical procedures (Bayes being solely and naïvely mentioned for classification, furthermore with data-based evaluation of the class “prior” probabilities, p211). At this stage, procedures are often only defined through spinets of code, like the private Theil-Sen estimator (pp204-205). The continuous case boils to a Normality assumption, with its pmf being defined (p212) as

\text{Pr}(x=\mu)=\frac{1}{\sqrt{2\pi\sigma}}e^{-(x-\mu)/2\sigma^2}

which contains at least three errors! Chapter 9 is the equivalent of Chapter 8 for machine learning, mostly centred on private gradient descent. And a Pytorch section (pp232-235). Completed by a light Chapter 10 on synthetic data, which does not seem to broach upon the issue of large dimension covariates, providing instead a list of GAN synthetizers.

Part III (Deploying differential privacy) is even more about practice, with Chapter 11 on privacy attacks, Chapter 12 on calibrating a privacy mechanism (co-written with Jayshree Sarathy), and good practice (like codebooks and data annotations), with the appearance of contextual integrity I discovered if not perfectly understood last year at the BIRS workshop in Kelowna. And Chapter 13 on planning a privacy project, with an 11 step checklist, most of which are quite vague [imho] and do include strategies to make the data owners confident their privacy is safe.

[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE]