Archive for data privacy

Bayesian, adversarial, oceanic, privacy

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , on March 6, 2026 by xi'an

We just arXived a new paper on Bayesian privacy! We meaning Cameron Bell, Antoine Luciano, Timothy Johnston and myself, as members of my ERC OCEAN lab at PariSanté and Paris Dauphine. While sharing the same ground as my recent paper with James Bailie, Joshua Bon and Judith Rousseau, this one is definitely more mainstream Bayesian in that the entire decision process falls under the Bayesian hat, with the ultimate decision being the choice of the release mechanism by the data holder (or hoarder!). To rationalise this decision process, we break the framework as resulting from the actions of three actors, namely the data holder, Alice, the data scientist, Bob, and the eavesdropper. Eve. (As in my earlier posts on solving Le Monde’s math puzzles, we could have used pronouns from other cultures, but I feared this would have confused some of the readers. Incidentally, I found out that the earliest use of the first two pronouns was within the groundbreaking cryptography 1977 paper of Rivest, Shamir and Adleman, bringing the RSA algorithm to the World! With Eve appearing in an early, highly-cited privacy paper by Montréal’s Bennett, Brassard, and (unconnected to me!) Robert, in 1988.)

We thus consider a Bayesian setting in which, given data x, held by Alice, inference is to be performed by Bob on a parameter θ. Performing such inference requires Alice releasing information derived from x, which may contain sensitive content, exploited by Eve. Our approach is to compare Alice’s release mechanisms according to both the quality of inference on θ (from Bob’s viewpoint) and the privacy leakage regarding x (sought by Eve and dreaded by Alice). To formalise this evaluation, we posit that Alice refers to a loss function that is a linear combination of Bob’s and Eve’s losses, the weight on Eve’s loss being then negative. (An alternative to be considered in future work is Alice using a ratio of Bob’s and Eve’s losses, possibly set to different powers, the rationale being that a zero loss for Eve is intolerable for Alice.) As in Bayesian experimental design, a prior on the data is necessary for Eve to infer on the hidden data based on the release mechanism and released output and for Alice to evaluate the risk of said release mechanism . (They may differ, as long as they are both made public.) To calibrate Alice’s loss, we opted for a balance that returns the same risk for a full data release and a total lack of release. In specific, informed, settings, other weights could be chosen. While finding the optimal release strategy is impossible but for highly discrete settings, the framework obviously allows for the ranking of natural strategies like insufficient statistics and synthetic datasets. Comments welcome!

statistics summer school at Warwick (July 2026)

Posted in Statistics, Travel, University life with tags , , , , , , , , , , , , , , , on November 26, 2025 by xi'an

Brussels snapshot [jatp]

Posted in pictures, Running, Travel with tags , , , , , , , , on May 31, 2025 by xi'an

JSM 2024, Portland, Day 1

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , on August 6, 2024 by xi'an

Strolling through the Oregon Conference Centre on 5 Aug, I am as always amazed at how the JSM conference centres have the ability to swallow in thousands of participants without giving an impression of overcrowding! (And appreciating the moderate air conditioning, which for once does not require wearing a fleece indoors!), I must admit that my first impressions of the city itself have been rather poor, as I was (unsuccessfully) seeking an after-hour grocery near the conference centre, I walked through run-down areas and kept passing homeless people, most in a sorry state. And hearing throughout the night, And again this morning in the warehouse maze I jogged through  before hitting the Willamette River path, which goes uninterrupted for miles. And showed me an unexpected spot, the Kevin Duckworth dock, where swimming the Willamette is feasible. Hopefully attempting a morn swim before I leave Portland.

I attended the quantum computing session, with a rather light introduction without bringing much light on the nature of qubits for storage and computing. In particular, the role of the complex coefficients of the 0 and 1 states. Then my friend Brani Vidakovic gave us an hand-on demonstration (almost hand-on as the conference facilities were unable to let a code run live!, eons away from quantum performances!!). Using Qiskit and Anaconda. He pointed out that measuring a qubit is destroying its quantum nature, the very equivalent of Schrôdinger’s cat. But does not make it clear whether or not the qubit later returns to a random entity, since frequency stabilisation is assumed, witness the histograms displayed by Brani. or a Nature paper of last year.

Went to a (poorly attended) privacy panel session next. With a defence of the value of differential privacy (Jordan Awan, Purdue) somewhat connected with our own (Ocean) work, although utility not understood in a decision-theoretic sense. And privacy remaining in an one-suits-all sense. With Michael Hawes from the US Census Bureau introducing more dimensions than mere DP (coarsening, suppression, swapping, &tc.), with legal aspects (Title 13) of disclosure risk. With the positive (for me) notion of providing a meaningful assessment of disclosure, with some records being more vulnerable than others. And another one on the cumulative disclosure risk over time. If short in quantitative entries. With Valbona Bejleri (USDA NASS, with a Census of their own) on cell suppressions and metrics for assessing disclosure, eg Shannon’s information entropy. And with Gary Howarth (Privacy Engineering Program, NIST). whose point remained rather unclear to me, like the apparently obvious point that adding features increase dispersion between populations. Although presenting tools for convincing experts and actors of the efficiency of privacy protection techniques.

Which continued (sort of) on the afternoon with a synthetic data for preserving privacy panel session. With Bradley Malin (Vanderbilt U) on the dangers of (Nature Communication paper of 2022). And Harrison Quick using a posterior predictive to achieve differential privacy, albeit considering the posterior predictive as the statistical analysis outcome does not seem the right focus (and recoup my earlier criticism of differential privacy requiring bending one’s prior beliefs). Indeed, as a Bayesian aiming at inference rather than merely at not releasing raw data, I would use synthetic data generated from that posterior predictive to return a posterior on the parameters of interest. And Joshua Snoke (RAND) with cautionary warnings. Like producing “invalid” inference because the reliance on a specific model (but isn’t that the case for most of statistics?). And Roee Gutman (Brown U), discussing the special case of record linkage. More into the difficulties in creating synthetic data. (Like the issue with missingness.) Somehow disappointing in not reaching a more statistical and quantitative perspective, eg by sticking to a Bayesian perspective the whole way.

On a non-technical side, I am surprised at hardly anyone adopting the bring-your-own-container policy in the OCC cafés, or outside, given the large number of attendees carrying one or several liquid containers.

data protection [not from Les Houches]

Posted in Books, Mountains, Statistics with tags , , , , , , , , , , , on March 16, 2024 by xi'an

While running a “kitchen” workshop on Bayesian privacy in Les Houches, Le Monde published on π day a recap of a recent report on AI commanded by the French Government. Among other things, it contains recommendations on alleviating the administrative blocks in accessing personal data, based on a model for data protection created decades earlier around the CNIL structure. The final paragraph wishes for the creation of a “laboratory” that would test collaborative, altruistic, efficient models towards sharing data for learning, which is one of the main goals of OCEAN. Without mentioning any technical aspect, like an adoption of some privacy measure at a national or European level.