Archive for Bayesian privacy

persuasive (and Oceanic) privacy

Posted in Books, Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on February 3, 2026 by xi'an

I am quite excited about the paper James Baillie, Joshua Bon, Judith Rousseau, and myself just arXived! A novel framework for measuring privacy we have been working on for at least the past year, partly through the previous Les Houches privacy workshops. In the spirit of these workshops and the larger scale ERC Synergy grant OCEAN, we develop therein a rather generic Bayesian game-theoretic perspective on achieving statistical privacy. It involves a Sender (observing the original data and delivering a limited output) and a Receiver (with potential adversarial intentions). The paper mostly focus on setting a theoretical framework, including the creation of new, purpose-driven privacy definitions that are rigorously justified, while also allowing for the assessment of existing privacy guarantees through game theory. While this was not our original intent, we show that pure and probabilistic differential privacy notions, in the Dwork et al. (2006) sense, are special cases of our framework. This setting provides new interpretations of the post-processing inequality. Furthermore, and somewhat more importantly, we also prove that our privacy guarantees can be established for deterministic algorithms, which are outside current privacy standards. Hopefully, we’ll make further progress at the incoming privacy workshop next month, to be held in Venice (again).

meanwhile, today, in Nashville

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , on August 5, 2025 by xi'an

[Josh and I organised this privacy session a while ago, before Josh returned to Adelaïde, and we found neither of us could/would attend JSM!]

Data Privacy: Frontiers and Barriers of Differential Privacy

 

Steven MacEachern Chair and Discussant
The Ohio State University

Joshua Bon & Christian Robert Organizers
Université Paris Dauphine, ERC OCEAN

Monday, Aug 4: 2:00 PM – 3:50 PM

Main Sponsor

Committee on Privacy and Confidentiality

Presentations

Enhancing Feature-Specific Data Protection via Bayesian Coordinate Differential Privacy

Local Differential Privacy (LDP) offers strong privacy guarantees without requiring users to trust external parties. However, LDP applies uniform protection to all data features, including less sensitive ones, which degrades performance of downstream tasks. To overcome this limitation, we propose a Bayesian framework, Bayesian Coordinate Differential Privacy (BCDP), that enables feature-specific privacy quantification. This more nuanced approach complements LDP by adjusting privacy protection according to the sensitivity of each feature, enabling improved performance of downstream tasks without compromising privacy. We characterize the properties of BCDP and articulate its connections with standard non-Bayesian privacy frameworks. We further apply our BCDP framework to the problems of private mean estimation and ordinary least-squares regression. The BCDP-based approach obtains improved accuracy compared to a purely LDP-based approach, without compromising on privacy.

Speaker: Alireza Fallah, UC Berkeley


Composition of privacy mechanisms: Only fresh noise counts

Composition is a key desiderata for a differential privacy (DP) flavor because it ensures a controlled degradation of the total privacy loss as additional statistics are released. However, ever since Pufferfish DP flavors were first introduced in 2012 it has remained an open problem whether these flavors – which take a Bayesian viewpoint by incorporating into DP the attacker’s uncertainty in the confidential data – satisfy composition. In this work, we resolve this question by proving that a Pufferfish flavor satisfies composition if and only if it is equivalent to a pure ε-DP flavor. Therefore, the generalization of DP to Pufferfish privacy is incompatible with the desiderata of composition. Furthermore, we determine that a Pufferfish mechanism composes with itself if and only if it satisfies pure ε-DP – i.e. if and only if it does not make use of the attacker’s uncertainty in the data generation mechanism. This result establishes that any class of composable mechanisms which satisfy Pufferfish – such as those found in existing literature – in fact satisfy the stronger condition of pure ε-DP. In intuitive terms, we show that a composable mechanism cannot reuse the “noise” provided by the attacker’s Bayesian model of the confidential data: only fresh noise counts when it comes to composition.

Speaker: James Bailie, Harvard University


Tukey Depth Mechanisms for Practical Private Mean Estimation

Mean estimation is a fundamental task in statistics and a focus within differentially private statistical estimation. While univariate methods based on the Gaussian mechanism are widely used in practice, more advanced techniques such as the exponential mechanism over quantiles offer robustness in the strong contamination model and improved performance, especially for small sample sizes. Tukey depth mechanisms carry these advantages to multivariate data, providing similar strong theoretical guarantees. However, practical implementations fall behind these theoretical developments.
In this talk, I will discuss first steps to bridge this gap by implementing the (Restricted) Tukey Depth Mechanism, a theoretically optimal mean estimator for multivariate Gaussian distributions, yielding improved practical methods for private mean estimation. The implementations enable the use of these mechanisms for small sample sizes or low-dimensional data. Additionally, I will present variants of these mechanisms that use approximate versions of Tukey depth, trading off accuracy for faster computation. We demonstrate their efficiency in practice, showing that they are viable options for modest dimensions. Given their strong accuracy and robustness guarantees, we contend that they are competitive approaches for mean estimation in this regime. Finally, I will discuss future directions for improving the computational efficiency of these algorithms by leveraging fast polytope volume approximation techniques, paving the way for more accurate private mean estimation in higher dimensions, as well as conjectured barriers toward this goal.

This talk is based on joint work with Gavin Brown.

Speaker: Lydia Zakynthinou, UC Berkeley

Bayesian decision-theory for data privacy [surfin’ the Oce’n, 30 April, INRIA Paris]

Posted in Statistics, University life with tags , , , , , , , , , , , , , , , , , on April 23, 2025 by xi'an

Abstract

The scientific and economic value of data continues to grow alongside technology advances. New hardware and software developments enable, but often require, larger and more complex datasets to function effectively. As the importance of input data to these systems becomes increasingly recognized, so too does the loss of privacy for data providers. In this context, data privacy emerges as a critical issue for fields such as statistics and machine learning, as well as for scientific and industrial endeavours that rely on sensitive data. We propose a framework for measuring privacy from a Bayesian decision-theoretic perspective. This framework enables the creation of new, purpose-driven privacy principles that are rigorously justified, while also allowing for the assessment of existing privacy definitions through decision theory. We pay particular attention to the privacy of deterministic algorithms, which are overlooked by current privacy standards, and to the privacy of N Monte Carlo samples drawn from an invariant distribution as N goes to infinity. We show that Probabilistic Differential Privacy is a special case of our framework and provide some new interpretations for Differential Privacy as a result.

handbook of sharing confidential data [book review]

Posted in Statistics with tags , , , , , , , , , , , , , on March 12, 2025 by xi'an

A new Chapman & Hall handbook appeared on the most current issue of confidentiality and privacy, which has been edited by Jörg Drechsler, Daniel Kifer, Jerome Reiter, and Aleksandra Slavković. The forty authors of the 18 chapters are mostly from the U.S., with a few outliers from Edinburgh (involved in two chapters on protecting the Scottish Longitudinal Study and the U.S. IRS tax data) and Tallinn (for a chapter on secure multi-party computation applications). This means a more U.S. centric focus for realistic implementations as, e.g., with the Census Bureau (which employs 25% of the authors), than those implied by EU regulations, for instance.

Overall, I enjoyed reading these chapters and would certainly use the book as a first entry to a graduate course on privacy (as opposed to some books I recently reviewed). The first two chapters are 100% formula-free and thus more surveys than informative entries to the field, imho. The following Part II on formal privacy techniques covers the expected standards of differential privacy, local vs. global design, single vs. multiple queries, consequence on learning machines and statistical procedures. Concerning Bayesian aspects, Chapter 7 about private machine learning has two paragraphs on the privacy properties of MCMC algorithms albeit not exposing clearly enough that privacy vanishes as the number of iterations grows to infinity. Chapter 8 concentrates on statistical differential privacy, much along my own perception of the requirements for a genuine statistical approach, with Bayesian aspects not sidelined. If less critical of differential privacy than I. Chapter 9 focusses on system issues, investing a dozen pages into the specifics of pseudo-random generators. Part III is about synthetic data, with some overlap between the first two chapters. (I would deem DP need not be introduced by Chapter 12.) I find the section rather superficial, mostly formula free, and lacking in the statistical impact.

As an aside, I am disappointed at the poor rendering of (mathematical) equations making me wonder which type of LaTeX, if any, was used. There are even genuine typos  that seem to result from cut and past encoding errors (see, e.g., the final accentuated c of Sklavković). The reference lists are plentiful, see e.g. the 164 entries for Chapter 7, to the point it would have made more sense to regroup them into a single bibliography. (The predictable reply being that chapters are sold separately and need their respective reference lists.)

[Disclaimer about potential self-plagiarism: this post or an edited version of it could possibly appear in my Books Review section in CHANCE.]