Archive for Bayesian privacy

no privacy in nature

Posted in Books, Statistics, University life with tags , , , , , , , , , on September 26, 2026 by xi'an

A paper about privacy (or lack thereof) in Nature by Knolle et al. about medical AI models that exhibit a weakness to membership privacy attacks thru multiple queries of the public model. As discussed in a commentary article by Zhang & Ghassemi, a membership privacy attack proves successful when the confidence of the model prediction jumps to higher values for a real target of interest compared with an imaginary one. This obviously assume that users (and attackers) may repeat queries ad nauseam from the model, which differs from our Bayesian privacy setting where the output is provided once and only once (whatever the release mechanism is).

The attacker is modelled as resorting to a basic likelihood-ratio MIAs3,4, that is, a test based on the prediction confidence attached to the target model with the null being that the target is not a member:

“the parameters of the distributions under the two hypotheses are specified by parametric fitting of sample confidence values obtained from reference models. Reference models are models assumed to be trained by the attacker and are ideally, but not necessarily, of similar architecture as the target model and trained on data similar to the training dataset” – M Knolle & al.

but the paper does not provide further details about inferring about the model parameters. The following sentence is also unclear

“objectively larger threats are posed by privacy attacks with stronger assumptions on a potential attacker, such as access to model parameters17, access to parameter updates during model training18 (…) By contrast, the type of attack we consider here requires querying the target model only once (to obtain a prediction for the target record)” – M Knolle & al.

in that, indeed, returning the model parameter estimates is more informative about the data than a black-box prediction interface, but I miss the single query point as it seems to me that the attacker must multiply the queries to build a confidence distribution under both hypotheses.

“Purely technical measures, such as a mathematical approach called differential privacy9, often create performance trade-offs10 that are too limiting for medical AI tools. Instead, what is needed are regulatory and sociotechnical safeguards, such as privacy audits and risk assessments, that are specific to the domain in question.” – H Zhang & M Ghassemi

“our results indicate that privacy attacks against AI models may be much more effective at compromising the privacy of individual data contributors than previously thought. This suggests that current AI privacy risk reporting practices may underestimate individual-level risk and thus motivates the integration of mathematically verifiable risk mitigation strategies such as differential privacy (DP) into medical AI model development workflows.” – M Knolle & al.

“the finding that record-level differential privacy is insufficient for multi-record patients is particularly impactful and has clear policy and implementation implications” –Referee #2

“we also found that the full mitigation of MIAs for all data-contributing patients requires stricter levels of privacy protection (ϵ, δ ) smaller than previously believed. Moreover, our results also show that fully mitigating MIAs requires DP accounting at the patient- rather than the record-level.” – M Knolle & al.

In a funny (?) clash, the authors of the paper and those of the comment and a referee seem to disagree on the pertinence of differential privacy guarantees in this context. But the conclusion remains very vague on a rigorous way to assert privacy leaks and confidentiality protection for a given AI and its supporting dataset.

 

 

Bayesian persuasive privacy at ICML²⁶

Posted in Books, Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , on May 14, 2026 by xi'an

Bayesian privacies [slides]

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , , , , on April 4, 2026 by xi'an

di ritorno a Venezia, nella privacy oceanica

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , on March 15, 2026 by xi'an

Bayesian, adversarial, oceanic, privacy

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , on March 6, 2026 by xi'an

We just arXived a new paper on Bayesian privacy! We meaning Cameron Bell, Antoine Luciano, Timothy Johnston and myself, as members of my ERC OCEAN lab at PariSanté and Paris Dauphine. While sharing the same ground as my recent paper with James Bailie, Joshua Bon and Judith Rousseau, this one is definitely more mainstream Bayesian in that the entire decision process falls under the Bayesian hat, with the ultimate decision being the choice of the release mechanism by the data holder (or hoarder!). To rationalise this decision process, we break the framework as resulting from the actions of three actors, namely the data holder, Alice, the data scientist, Bob, and the eavesdropper. Eve. (As in my earlier posts on solving Le Monde’s math puzzles, we could have used pronouns from other cultures, but I feared this would have confused some of the readers. Incidentally, I found out that the earliest use of the first two pronouns was within the groundbreaking cryptography 1977 paper of Rivest, Shamir and Adleman, bringing the RSA algorithm to the World! With Eve appearing in an early, highly-cited privacy paper by Montréal’s Bennett, Brassard, and (unconnected to me!) Robert, in 1988.)

We thus consider a Bayesian setting in which, given data x, held by Alice, inference is to be performed by Bob on a parameter θ. Performing such inference requires Alice releasing information derived from x, which may contain sensitive content, exploited by Eve. Our approach is to compare Alice’s release mechanisms according to both the quality of inference on θ (from Bob’s viewpoint) and the privacy leakage regarding x (sought by Eve and dreaded by Alice). To formalise this evaluation, we posit that Alice refers to a loss function that is a linear combination of Bob’s and Eve’s losses, the weight on Eve’s loss being then negative. (An alternative to be considered in future work is Alice using a ratio of Bob’s and Eve’s losses, possibly set to different powers, the rationale being that a zero loss for Eve is intolerable for Alice.) As in Bayesian experimental design, a prior on the data is necessary for Eve to infer on the hidden data based on the release mechanism and released output and for Alice to evaluate the risk of said release mechanism . (They may differ, as long as they are both made public.) To calibrate Alice’s loss, we opted for a balance that returns the same risk for a full data release and a total lack of release. In specific, informed, settings, other weights could be chosen. While finding the optimal release strategy is impossible but for highly discrete settings, the framework obviously allows for the ranking of natural strategies like insufficient statistics and synthetic datasets. Comments welcome!