Archive for Oregon

JSM 2024, Portland, Day 2

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , on August 7, 2024 by xi'an

By happenstance, I started my day in the cybersecurity session. With (again) hardly a soul in the room… A first talk on avoiding herding and achieving asymptotic truth learning (about a binary outcome) in a graph by putting constraints on the graph structure, without any clear connection with statistics or cybersecurity. Even less for the second talk on optimising masks. Only with the third one came cybersecurity motivations, the focus being on a two-player Stackelberg game already used in this framework. The result proper was about estimating the parameter of a (rather unrealistic) posited model reproducing the adversarial actions. The last talk about jailbreak attacks was again off-field by miles.


Then attended (as intended!) the 2024 Blackwell-Rosenbluth Award session, featuring the nominees Sharmistha Guha (Texas A&M) on multiple network inference, Simon Mak (Duke) on using Bayesian surrogate models, Guanyang Wang (Rutgers) who recently spoke at our mostly Monte Carlo seminar, Akihiko Nishimura (John Hopkins) on a unification of HMC and PDMPs, most appropriate when located next to Mount Hamilton and the Zigzag river!, and Maria Skoularidou (MIT) on evolutionary Monte Carlo for gene expression. Alas with hardly anyone in the room.

In the afternoon, I went to the (very well-attended this time!, with no seat available for many attendees, incl. yours truly!) COPSS Elizabeth L. Scott Lecture by my friend from Rutgers, Regina Liu, on the highly relevant challenge of combining inferences from diverse data sources. Using the (definitely Rutgerian!) approach of confidence distributions!

Last (late) afternoon, I went swimming from Kevin Duckworth dock, just below the conference centre. Water was quite warm (and green), with a few other swimmers, and no stomachical after-effect, so far. Hence I returned there once again this afternoon.

JSM 2024, Portland, Day 1

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , on August 6, 2024 by xi'an

Strolling through the Oregon Conference Centre on 5 Aug, I am as always amazed at how the JSM conference centres have the ability to swallow in thousands of participants without giving an impression of overcrowding! (And appreciating the moderate air conditioning, which for once does not require wearing a fleece indoors!), I must admit that my first impressions of the city itself have been rather poor, as I was (unsuccessfully) seeking an after-hour grocery near the conference centre, I walked through run-down areas and kept passing homeless people, most in a sorry state. And hearing throughout the night, And again this morning in the warehouse maze I jogged through  before hitting the Willamette River path, which goes uninterrupted for miles. And showed me an unexpected spot, the Kevin Duckworth dock, where swimming the Willamette is feasible. Hopefully attempting a morn swim before I leave Portland.

I attended the quantum computing session, with a rather light introduction without bringing much light on the nature of qubits for storage and computing. In particular, the role of the complex coefficients of the 0 and 1 states. Then my friend Brani Vidakovic gave us an hand-on demonstration (almost hand-on as the conference facilities were unable to let a code run live!, eons away from quantum performances!!). Using Qiskit and Anaconda. He pointed out that measuring a qubit is destroying its quantum nature, the very equivalent of Schrôdinger’s cat. But does not make it clear whether or not the qubit later returns to a random entity, since frequency stabilisation is assumed, witness the histograms displayed by Brani. or a Nature paper of last year.

Went to a (poorly attended) privacy panel session next. With a defence of the value of differential privacy (Jordan Awan, Purdue) somewhat connected with our own (Ocean) work, although utility not understood in a decision-theoretic sense. And privacy remaining in an one-suits-all sense. With Michael Hawes from the US Census Bureau introducing more dimensions than mere DP (coarsening, suppression, swapping, &tc.), with legal aspects (Title 13) of disclosure risk. With the positive (for me) notion of providing a meaningful assessment of disclosure, with some records being more vulnerable than others. And another one on the cumulative disclosure risk over time. If short in quantitative entries. With Valbona Bejleri (USDA NASS, with a Census of their own) on cell suppressions and metrics for assessing disclosure, eg Shannon’s information entropy. And with Gary Howarth (Privacy Engineering Program, NIST). whose point remained rather unclear to me, like the apparently obvious point that adding features increase dispersion between populations. Although presenting tools for convincing experts and actors of the efficiency of privacy protection techniques.

Which continued (sort of) on the afternoon with a synthetic data for preserving privacy panel session. With Bradley Malin (Vanderbilt U) on the dangers of (Nature Communication paper of 2022). And Harrison Quick using a posterior predictive to achieve differential privacy, albeit considering the posterior predictive as the statistical analysis outcome does not seem the right focus (and recoup my earlier criticism of differential privacy requiring bending one’s prior beliefs). Indeed, as a Bayesian aiming at inference rather than merely at not releasing raw data, I would use synthetic data generated from that posterior predictive to return a posterior on the parameters of interest. And Joshua Snoke (RAND) with cautionary warnings. Like producing “invalid” inference because the reliance on a specific model (but isn’t that the case for most of statistics?). And Roee Gutman (Brown U), discussing the special case of record linkage. More into the difficulties in creating synthetic data. (Like the issue with missingness.) Somehow disappointing in not reaching a more statistical and quantitative perspective, eg by sticking to a Bayesian perspective the whole way.

On a non-technical side, I am surprised at hardly anyone adopting the bring-your-own-container policy in the OCC cafés, or outside, given the large number of attendees carrying one or several liquid containers.

Mount Hood [jatp]

Posted in Mountains, pictures, Running with tags , , , , , , , , , , , , , , , , , on August 5, 2024 by xi'an

JSM 2024, Portland, OR, is off!

Posted in pictures, Statistics, Travel, University life with tags , , , , , , , , on August 4, 2024 by xi'an

differential privacy for Bayesian inference

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , on July 30, 2024 by xi'an

As I was reading it in preparation for my JSM²⁴ lecture, I found anew that, in this landmark paper of Dimitrikakis et al. (2017), some limitations of the concept of differential privacy were most apparent:

– a requirement to bend both the model and the prior to fit differential privacy, like switching to Lipschitz constraints or using new (e.g., truncated) priors, which runs contrary to Bayesian principles, although the former can be seen as a form of randomization akin to ABC when the randomization itself is accounted for in the derivation of the “exact” posterior distribution (as in the paper of Berah, Favaro, and Rao (2023) on running MCMC for Bayesian non-parametric estimation on privatized (noisy) data I discussed a few days ago);

– a subtle switch of the randomness from the (privatization) procedure itself (as in Dwork (2006)) to the uncertainty about the parameter, not that it clashes per se with Bayesian principles (even though there is an unclear randomness statement in Theorem 9, when the prior itself seems to become random (?)). Which actually means that producing one realisation from the posterior is the (privatization) procedure, as I realised when discussing with Shenggang Hu in Warwick;

– a linear degradation of the privacy parameter ε when moving from one realisation of the posterior to a simulated sample, assuming iid realisations (I wonder whether or not releasing a dependent sample could involve the ESS instead of the number of MCMC iterations). The upper bound means that no privacy whatsoever is guaranteed for an infinite posterior sample, hence for delivering de facto the posterior (despite the paper producing an (ε,0) bound on the Kullback-Leibler measure of the difference between posteriors);

– an absence of prior knowledge or modelling on the data itself, unless the distance from the actual data to an hypothetical alternative, ρ(x,y), can be interpreted as minus a score function conditional on the actual data, eg the opposite of the log predictive

  • the occurrence of an “exponential prior” that is exp p/m ell (theta) with ell a Lipschitz constant for the associated likelihood

  • the assumption that the data user is not an adversary of the data keeper, in that their utility function is about the parameter (and publicly available)