Archive for synthetic data

Bayesian Adversarial Privacy [v2]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , on September 11, 2026 by xi'an

We have just reposted our paper Bayesian Adversarial Privacy on arXiv to reflect the revision we wrote in the past months, to address the (quite sensible) comments from the reviewers. Interestingly the discussants of my Akaike lecture made similar points. The main changes are in explaining more clearly the nature of the combined loss, with Antoine coming up with the use of illuminating R-U map representations, in enlarging the references to other approaches, in stressing that Eve was an Alice’s construct rather than a genuine adversary, but still integrating the case of “multiple Eves”, in mellowing our criticisms of DP, and in expanding the conclusion with limitations and extensions subsections.

Bayesian Privacy [Akaike Lecture slides]

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , on September 10, 2026 by xi'an

JSM 2024, Portland, Day 1

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , on August 6, 2024 by xi'an

Strolling through the Oregon Conference Centre on 5 Aug, I am as always amazed at how the JSM conference centres have the ability to swallow in thousands of participants without giving an impression of overcrowding! (And appreciating the moderate air conditioning, which for once does not require wearing a fleece indoors!), I must admit that my first impressions of the city itself have been rather poor, as I was (unsuccessfully) seeking an after-hour grocery near the conference centre, I walked through run-down areas and kept passing homeless people, most in a sorry state. And hearing throughout the night, And again this morning in the warehouse maze I jogged through  before hitting the Willamette River path, which goes uninterrupted for miles. And showed me an unexpected spot, the Kevin Duckworth dock, where swimming the Willamette is feasible. Hopefully attempting a morn swim before I leave Portland.

I attended the quantum computing session, with a rather light introduction without bringing much light on the nature of qubits for storage and computing. In particular, the role of the complex coefficients of the 0 and 1 states. Then my friend Brani Vidakovic gave us an hand-on demonstration (almost hand-on as the conference facilities were unable to let a code run live!, eons away from quantum performances!!). Using Qiskit and Anaconda. He pointed out that measuring a qubit is destroying its quantum nature, the very equivalent of Schrôdinger’s cat. But does not make it clear whether or not the qubit later returns to a random entity, since frequency stabilisation is assumed, witness the histograms displayed by Brani. or a Nature paper of last year.

Went to a (poorly attended) privacy panel session next. With a defence of the value of differential privacy (Jordan Awan, Purdue) somewhat connected with our own (Ocean) work, although utility not understood in a decision-theoretic sense. And privacy remaining in an one-suits-all sense. With Michael Hawes from the US Census Bureau introducing more dimensions than mere DP (coarsening, suppression, swapping, &tc.), with legal aspects (Title 13) of disclosure risk. With the positive (for me) notion of providing a meaningful assessment of disclosure, with some records being more vulnerable than others. And another one on the cumulative disclosure risk over time. If short in quantitative entries. With Valbona Bejleri (USDA NASS, with a Census of their own) on cell suppressions and metrics for assessing disclosure, eg Shannon’s information entropy. And with Gary Howarth (Privacy Engineering Program, NIST). whose point remained rather unclear to me, like the apparently obvious point that adding features increase dispersion between populations. Although presenting tools for convincing experts and actors of the efficiency of privacy protection techniques.

Which continued (sort of) on the afternoon with a synthetic data for preserving privacy panel session. With Bradley Malin (Vanderbilt U) on the dangers of (Nature Communication paper of 2022). And Harrison Quick using a posterior predictive to achieve differential privacy, albeit considering the posterior predictive as the statistical analysis outcome does not seem the right focus (and recoup my earlier criticism of differential privacy requiring bending one’s prior beliefs). Indeed, as a Bayesian aiming at inference rather than merely at not releasing raw data, I would use synthetic data generated from that posterior predictive to return a posterior on the parameters of interest. And Joshua Snoke (RAND) with cautionary warnings. Like producing “invalid” inference because the reliance on a specific model (but isn’t that the case for most of statistics?). And Roee Gutman (Brown U), discussing the special case of record linkage. More into the difficulties in creating synthetic data. (Like the issue with missingness.) Somehow disappointing in not reaching a more statistical and quantitative perspective, eg by sticking to a Bayesian perspective the whole way.

On a non-technical side, I am surprised at hardly anyone adopting the bring-your-own-container policy in the OCC cafés, or outside, given the large number of attendees carrying one or several liquid containers.

[strong] foundations of synthetic B’earning

Posted in Books, Statistics, University life with tags , , , , , , , , , , on July 15, 2024 by xi'an

I only recently read the (foundational!) paper on Foundations of Bayesian learnin from synthetic data by Harrison Wilde, Jack Jewson, Sebastian Volmer [all associated with Warwick at some point] and [my long time friend]  Chris Holmes, that merges Bayesian inference with differential privacy constraints thru generalised / Gibbs interface. Recouping with the M-open perspective in order to accommodate the misspecified nature of synthetic data. I like the approach very much in that it intersects a lot with my own views, excepts for following the differential privacy formalism. I however think that further progress could be made by adopting an even more Bayesian position.

Their key messages from that paper are that

  1. learning from synthetic data may prove damaging to your (data) health
  2. robustness unsurprisingly reduces the odds or magnitude of the damage
  3. real data can still be used to some extent

Since the (synthetic) generating model can be a GAN, the privacy requirement is such that noise is “injected” in the input data and in the learning mechanism. This is not discussed in the paper but highly conservative constraints surely make the DGP loose several learning points.

On the side, I also like the alternative of opposing data keeper and learner rather than data owner and adversary. Learning here means taking an optimal B decision about the actual data averaged over the true DGP. With a prior on the distribution of the actual data, while being unable to avoid misspecification in representing the (marginal) synthetic generation model.

Unsurprisingly, the alternative approach is relying on proper scoring rules as Bissiri et al. (2016). Rather than finding the distribution KL closest to the synthetic generative model, robustified by generalised B inference, either via downweighting or via ß-divergence. With a preference for the latter. Interestingly, the authors consider the optimal learning size for the synthetic data. Since bringing in more synthetic data does not mean better performances.

Contextual Integrity for Differential Privacy #4 [23w5106]

Posted in Books, Mountains, pictures, Running, Statistics, Travel, University life, Wines with tags , , , , , , , , , , , , , , , , , , , , , , , , , , on August 5, 2023 by xi'an

Mostly short talks. First talk by Thomas Seinke (Google) on interpreting ε, with a side wondering of mine on the relation between exp(ε) and the uncertainty that comes with Monte Carlo outcome. Which may relate to this 2022 paper by Ruobin Gong. Second talk by Gautam Kamath (U Waterloo) on large language models under privacy with “public” data. Questioning the appropriateness of ML benchmarks in terms of privacy. Third talk by Mark Bun (Boston U) on replicability, privacy and adaptive generalisation in machine learning, with a strange criticism of confidence intervals on the same parameter not intersecting for two independent studies. And proposing high probability replicable algorithms that can be put in duality with differentially private algorithms at the cost of lowering precision and effective sample size. We also had another group discussion on how to reach out about privacy guarantees, which made me realise there were GDPR compliance software available.

In the afternoon session, Shlomi Hod (Boston U) presented a practical case of designing a privacy preserving protocol for the Israeli birth record. With a strong opposition from stakeholders to use synthetic data, due to a semantic drift from synthetic to manipulated to fake, to lying. Wanrong Zhang did not talk about her stunning recent ICML paper but instead of another practical case connected with mobile based Covid case predictions, by adding minimal noise to mobility data. Nidhi Hegde (U Alberta) gave up talking on Thomson sampling with privacy protection, to focus on an ongoing health application for Alberta as more suited for the workshop. And Ria Safavi-Naini (U Calgary) drew a parallel between information theory and DP versus CI.

While the workshop was scheduled till Friday noon, in usual BIRS habits (!), the morning session was cancelled for most people leaving Kelowna in the morning.