Archive for disclosure risk

Trump goes to war against… randomization!

Posted in Kids, Statistics, University life with tags , , , , , , , , , , on July 11, 2026 by xi'an

The US Department of Commerce, through its Office of Privacy and Open Government, has posted an order for federal statistical agencies to stop using randomization as a (worthwhile!) mean of protecting confidential data. Another instance of going against scientific arguments and validated procedures, and of dismantling decade long expertises, and of stopping a long series of observations, and of destasbilising federal agencies, from the Trump administration. To wit:

b.     Statistical Policy Directive No. 1, codified at 44 U.S.C. § 3563, requires each statistical agency, as enabled by the head of the parent agency, to adopt policies, best practices, and appropriate procedures to ensure the confidentiality, accuracy, and objectivity of its statistical activities and to disseminate timely and relevant statistical products.

SECTION 5. POLICY.

.01    Ensuring the Accuracy, Confidentiality, Objectivity, and Relevance of Statistical Products

a.     The Department shall, as a primary objective, aim to fulfill its statistical obligations by providing the public with accurate and objective information.

b.     The Department is firmly committed to striking a balance of accuracy, confidentiality, objectivity, and relevance for each statistical product that is consistent with its statistical obligations and the applicable legal requirements.

c.     Any use of noise infusion is inconsistent with the Department’s policies.

.02   The Census Bureau and the Bureau of Economic Analysis shall adhere to the following order of priority when considering and applying Disclosure Avoidance:

a.     Coarsening shall be the preferred category of Disclosure Avoidance methods for all statistical products.

b.     Suppression shall be permitted as a last resort, only to be used when coarsening is prohibited by law or would substantially defeat the accuracy or usability of a statistical product.

c.     Noise infusion shall not be used for any statistical product.

.03   This policy shall not be interpreted to conflict with any constitutional, statutory, regulatory, or other legal provision.

So here go (to the drain) differential privacy, enforced by the Census Bureau under the direction of John Abowd, as well as synthetic data, subsampling, et cetera! This will directly impact the coming 2030 census. With the asinine arguments by Commerce Department Kristen Eichamer that this is done to “maintain public confidence in our data while upholding our duty to safeguard the privacy of those who provide information” and that “indiscriminate use of noise infusion—even when not mandated by law—ultimately undermined confidence in the department’s products and cast doubt on their integrity.” NPR relates this order to the introduction of differentially private measures around the 2020 census and irrational Republican attacks on the procedure. Which reminds me of an earlier opposition to random sampling instead of “exhaustive enumeration” to correct for undercounts, which brought statisticians from both sides almost to blows! (With a defence of the correction procedure by my later friend Steve Fienberg.)

Bayesian, adversarial, oceanic, privacy

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , on March 6, 2026 by xi'an

We just arXived a new paper on Bayesian privacy! We meaning Cameron Bell, Antoine Luciano, Timothy Johnston and myself, as members of my ERC OCEAN lab at PariSanté and Paris Dauphine. While sharing the same ground as my recent paper with James Bailie, Joshua Bon and Judith Rousseau, this one is definitely more mainstream Bayesian in that the entire decision process falls under the Bayesian hat, with the ultimate decision being the choice of the release mechanism by the data holder (or hoarder!). To rationalise this decision process, we break the framework as resulting from the actions of three actors, namely the data holder, Alice, the data scientist, Bob, and the eavesdropper. Eve. (As in my earlier posts on solving Le Monde’s math puzzles, we could have used pronouns from other cultures, but I feared this would have confused some of the readers. Incidentally, I found out that the earliest use of the first two pronouns was within the groundbreaking cryptography 1977 paper of Rivest, Shamir and Adleman, bringing the RSA algorithm to the World! With Eve appearing in an early, highly-cited privacy paper by Montréal’s Bennett, Brassard, and (unconnected to me!) Robert, in 1988.)

We thus consider a Bayesian setting in which, given data x, held by Alice, inference is to be performed by Bob on a parameter θ. Performing such inference requires Alice releasing information derived from x, which may contain sensitive content, exploited by Eve. Our approach is to compare Alice’s release mechanisms according to both the quality of inference on θ (from Bob’s viewpoint) and the privacy leakage regarding x (sought by Eve and dreaded by Alice). To formalise this evaluation, we posit that Alice refers to a loss function that is a linear combination of Bob’s and Eve’s losses, the weight on Eve’s loss being then negative. (An alternative to be considered in future work is Alice using a ratio of Bob’s and Eve’s losses, possibly set to different powers, the rationale being that a zero loss for Eve is intolerable for Alice.) As in Bayesian experimental design, a prior on the data is necessary for Eve to infer on the hidden data based on the release mechanism and released output and for Alice to evaluate the risk of said release mechanism . (They may differ, as long as they are both made public.) To calibrate Alice’s loss, we opted for a balance that returns the same risk for a full data release and a total lack of release. In specific, informed, settings, other weights could be chosen. While finding the optimal release strategy is impossible but for highly discrete settings, the framework obviously allows for the ranking of natural strategies like insufficient statistics and synthetic datasets. Comments welcome!

JSM 2024, Portland, Day 1

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , on August 6, 2024 by xi'an

Strolling through the Oregon Conference Centre on 5 Aug, I am as always amazed at how the JSM conference centres have the ability to swallow in thousands of participants without giving an impression of overcrowding! (And appreciating the moderate air conditioning, which for once does not require wearing a fleece indoors!), I must admit that my first impressions of the city itself have been rather poor, as I was (unsuccessfully) seeking an after-hour grocery near the conference centre, I walked through run-down areas and kept passing homeless people, most in a sorry state. And hearing throughout the night, And again this morning in the warehouse maze I jogged through  before hitting the Willamette River path, which goes uninterrupted for miles. And showed me an unexpected spot, the Kevin Duckworth dock, where swimming the Willamette is feasible. Hopefully attempting a morn swim before I leave Portland.

I attended the quantum computing session, with a rather light introduction without bringing much light on the nature of qubits for storage and computing. In particular, the role of the complex coefficients of the 0 and 1 states. Then my friend Brani Vidakovic gave us an hand-on demonstration (almost hand-on as the conference facilities were unable to let a code run live!, eons away from quantum performances!!). Using Qiskit and Anaconda. He pointed out that measuring a qubit is destroying its quantum nature, the very equivalent of Schrôdinger’s cat. But does not make it clear whether or not the qubit later returns to a random entity, since frequency stabilisation is assumed, witness the histograms displayed by Brani. or a Nature paper of last year.

Went to a (poorly attended) privacy panel session next. With a defence of the value of differential privacy (Jordan Awan, Purdue) somewhat connected with our own (Ocean) work, although utility not understood in a decision-theoretic sense. And privacy remaining in an one-suits-all sense. With Michael Hawes from the US Census Bureau introducing more dimensions than mere DP (coarsening, suppression, swapping, &tc.), with legal aspects (Title 13) of disclosure risk. With the positive (for me) notion of providing a meaningful assessment of disclosure, with some records being more vulnerable than others. And another one on the cumulative disclosure risk over time. If short in quantitative entries. With Valbona Bejleri (USDA NASS, with a Census of their own) on cell suppressions and metrics for assessing disclosure, eg Shannon’s information entropy. And with Gary Howarth (Privacy Engineering Program, NIST). whose point remained rather unclear to me, like the apparently obvious point that adding features increase dispersion between populations. Although presenting tools for convincing experts and actors of the efficiency of privacy protection techniques.

Which continued (sort of) on the afternoon with a synthetic data for preserving privacy panel session. With Bradley Malin (Vanderbilt U) on the dangers of (Nature Communication paper of 2022). And Harrison Quick using a posterior predictive to achieve differential privacy, albeit considering the posterior predictive as the statistical analysis outcome does not seem the right focus (and recoup my earlier criticism of differential privacy requiring bending one’s prior beliefs). Indeed, as a Bayesian aiming at inference rather than merely at not releasing raw data, I would use synthetic data generated from that posterior predictive to return a posterior on the parameters of interest. And Joshua Snoke (RAND) with cautionary warnings. Like producing “invalid” inference because the reliance on a specific model (but isn’t that the case for most of statistics?). And Roee Gutman (Brown U), discussing the special case of record linkage. More into the difficulties in creating synthetic data. (Like the issue with missingness.) Somehow disappointing in not reaching a more statistical and quantitative perspective, eg by sticking to a Bayesian perspective the whole way.

On a non-technical side, I am surprised at hardly anyone adopting the bring-your-own-container policy in the OCC cafés, or outside, given the large number of attendees carrying one or several liquid containers.

CANSSI HSC seminar series

Posted in Statistics, University life with tags , , , , , , , on January 19, 2022 by xi'an