Archive for differential privacy

Nature tidbits [6 Aug 2026]

Posted in Books, pictures, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , on October 7, 2026 by xi'an

The 6 August issue of Nature arrived after my return from Japan, but it took me a while to catch up, the more because its cover story, Bloom service, about fertilising the oceans with iron as a carbon sink did not appeal much to the geoengineering sceptic in me. Nice pun though. Flotsam of interest for me:

  • an article on GPT-4 (and open-weight models) predicting the outcomes of 70 preregistered US survey experiments, 469 effects and 119,330 participants in total. The predicted treatment effects are highly correlated with the observed ones, about as well as pooled human forecasters do, including studies published after the training cutoff. But the predicted effect sizes are systematically too large, which is not a minor detail for a field still recovering from its replication crisis! Plus a survey of 460 social scientists about using LLMs for pilot testing, and a web app to “forecast” your own experiment’s effect before running it. (I wonder what a prior elicited this way would look like.)
  • a paper on membership inference attacks on medical AI models, showing that patients who differ from the majority are the easiest to identify as part of the training data, which I already discussed here. This fits the long-known issue with differential privacy and outliers.
  • a story on the four 2026 Fields Medals, announced at the ICM in Philadelphia, in fields far, far away:
    • Yu Deng (Chicago), for his work on PDEs, including a rigorous derivation of the Boltzmann equation from hard-sphere dynamics (with Zaher Hani and Xiao Ma), wave kinetic equations, and probabilistic approaches to nonlinear Schrödinger equations.
    • John Pardon (Stony Brook), for symplectic topology and enumerative geometry. He started as an undergraduate by answering a question of Gromov on knot distortion.
    • Jacob Tsimerman (Toronto), for extending techniques within arithmetic and complex algebraic geometry and solving conjectures in this field.
    • Hong Wang (NYU and IHES in Bures-sur-Yvette, almost a colleague!), for Fourier restriction, progress on Falconer’s conjecture and, with Joshua Zahl, the resolution of the Kakeya conjecture in three dimensions.
  • a feature on the scientific papers most cited in patents.
  • a news item on a Science paper by Blasi, Hamilton, Gray and Bowern, estimating that humanity may have spoken tens of thousands of languages between 3,000 and 1,000 years ago. The paper suggests a shift away from nomadic life triggered this diversity before it collapsed. I wonder at how deep the statistical linguistics inferring this rich past from the surviving sample goes. In a sense, the finding is not that surprising, given how isolated small prehistoric communities were, and how quickly a language drifts apart without contact. (I would also like to check how much the estimate depends on the assumed birth-and-death rates of languages, since the lost ones leave no trace to calibrate against…)
  • an editorial on the White House’s “new golden age for science” plan. It argues that such a golden age needs open borders and funding for the social sciences, public health and the humanities, not only AI and engineering. Alas, the now usual Trump.2.0 fare, with an AI funding roll-out in the news pages.
  • an emergency call by WHO director-general Tedros Adhanom Ghebreyesus on why the WHO has never been more needed, and news that the first volunteer received an Ebola vaccine three months into the current outbreak.
  • a forecast that the coming El Niño should be the largest (by a mind-blowing margin) on record, pushing 2027 temperatures to new highs. And already having severe impact on tropical weather.
  • a scary comment on the risks of nuclear-powered merchant ships, in the very issue dated on the anniversary of Hiroshima…

no privacy in nature

Posted in Books, Statistics, University life with tags , , , , , , , , , on September 26, 2026 by xi'an

A paper about privacy (or lack thereof) in Nature by Knolle et al. about medical AI models that exhibit a weakness to membership privacy attacks thru multiple queries of the public model. As discussed in a commentary article by Zhang & Ghassemi, a membership privacy attack proves successful when the confidence of the model prediction jumps to higher values for a real target of interest compared with an imaginary one. This obviously assume that users (and attackers) may repeat queries ad nauseam from the model, which differs from our Bayesian privacy setting where the output is provided once and only once (whatever the release mechanism is).

The attacker is modelled as resorting to a basic likelihood-ratio MIAs3,4, that is, a test based on the prediction confidence attached to the target model with the null being that the target is not a member:

“the parameters of the distributions under the two hypotheses are specified by parametric fitting of sample confidence values obtained from reference models. Reference models are models assumed to be trained by the attacker and are ideally, but not necessarily, of similar architecture as the target model and trained on data similar to the training dataset” – M Knolle & al.

but the paper does not provide further details about inferring about the model parameters. The following sentence is also unclear

“objectively larger threats are posed by privacy attacks with stronger assumptions on a potential attacker, such as access to model parameters17, access to parameter updates during model training18 (…) By contrast, the type of attack we consider here requires querying the target model only once (to obtain a prediction for the target record)” – M Knolle & al.

in that, indeed, returning the model parameter estimates is more informative about the data than a black-box prediction interface, but I miss the single query point as it seems to me that the attacker must multiply the queries to build a confidence distribution under both hypotheses.

“Purely technical measures, such as a mathematical approach called differential privacy9, often create performance trade-offs10 that are too limiting for medical AI tools. Instead, what is needed are regulatory and sociotechnical safeguards, such as privacy audits and risk assessments, that are specific to the domain in question.” – H Zhang & M Ghassemi

“our results indicate that privacy attacks against AI models may be much more effective at compromising the privacy of individual data contributors than previously thought. This suggests that current AI privacy risk reporting practices may underestimate individual-level risk and thus motivates the integration of mathematically verifiable risk mitigation strategies such as differential privacy (DP) into medical AI model development workflows.” – M Knolle & al.

“the finding that record-level differential privacy is insufficient for multi-record patients is particularly impactful and has clear policy and implementation implications” –Referee #2

“we also found that the full mitigation of MIAs for all data-contributing patients requires stricter levels of privacy protection (ϵ, δ ) smaller than previously believed. Moreover, our results also show that fully mitigating MIAs requires DP accounting at the patient- rather than the record-level.” – M Knolle & al.

In a funny (?) clash, the authors of the paper and those of the comment and a referee seem to disagree on the pertinence of differential privacy guarantees in this context. But the conclusion remains very vague on a rigorous way to assert privacy leaks and confidentiality protection for a given AI and its supporting dataset.

 

 

Bayesian Adversarial Privacy [v2]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , on September 11, 2026 by xi'an

We have just reposted our paper Bayesian Adversarial Privacy on arXiv to reflect the revision we wrote in the past months, to address the (quite sensible) comments from the reviewers. Interestingly the discussants of my Akaike lecture made similar points. The main changes are in explaining more clearly the nature of the combined loss, with Antoine coming up with the use of illuminating R-U map representations, in enlarging the references to other approaches, in stressing that Eve was an Alice’s construct rather than a genuine adversary, but still integrating the case of “multiple Eves”, in mellowing our criticisms of DP, and in expanding the conclusion with limitations and extensions subsections.

Bayesian Privacy [Akaike Lecture slides]

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , on September 10, 2026 by xi'an

persuasive privacy at ICML 2026

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , , , , , on July 7, 2026 by xi'an