
Archive for Joint Statistical Meeting
Portland murals [jatp]
Posted in Books, pictures, Running, Statistics, Travel, University life with tags airbnb, American Statistical Association, ASA, homless people, jatp, Joint Statistical Meeting, JSM 2024, murals, Oregon, Pacific North West, Portland, United States of America, Willamette River on September 3, 2024 by xi'an
JSM 2024, Portland, miniday 4
Posted in Books, pictures, Running, Statistics, Travel, University life with tags AI, American Statistical Association, ASA, Bayesian decision theory, Bayesian nonparametrics, Bayesian privacy, BIRS, British Columbia, classification, contextual integrity, data science, differential privacy, draw bridge, ERC, federated learning, generalised Bayesian inference, Joint Statistical Meeting, JSM 2024, Kelowna, Ocean, open water swimming, Oregon, Pacific North West, Portland, Seattle, United States of America, University of Washington, Willamette River on August 11, 2024 by xi'an
Final (half)day at JSM is always a sad thing as most people have left, people are busy dismantling booths and packing boxes, and the few remaining participants are fidgety and sitting on their suitcases (there may have been more suitcases than people in the main area that day!), cafés are minimally staffed or simply closed. Hence not the best time-slot to deliver one’s talk! Still, a few dozen people attended our session. The session topic was Bridge the Gap: Differential Privacy and Statistical Analysis, organised by Bei Jiang whom I met last summer in Kelowna, at a BIRS workshop on privacy. Where I spoke on setting up a complete decision-theoretic framework, as developed (and still in development) within our Ocean group, esp. Joshua Bon, Stan du Ché, and Judith Rousseau. (Rather than on the original plan of talking about convergence versus privacy, as a criticism of differential privacy.) The other talks were by Shurong Li, strongly set with differential privacy when record linkage is present, and Xuan Bi on a fully decentralised federated learning with local exchanges of global gradients that limit privacy leaks. (With a mention of gossip learning I hadn’t seen previously!)

I enjoyed even more the session due to Naisyin Wang giving a discussion on the three talks, in closes the particular because I had not seen her in years, if a few times since she was my teaching assistant in Cornell in a Bayesian decision theory class I was building on the spot (with Linda Zhao as a student!). Which nicely closes the loop given the topic of my talk, of which she was quite supportive! While calling for debiasing post-processing and wondering about a two-dimensional decision theoretic perspective rather than a unidimensional one under hard privacy constraints.

On the way out, after parting from the few friends remaining in the convention centre, I spotted the nearby (steel) bridge being raised, although I could not see the boat responsible for it. I had been unaware of this possibility while running over and under it, as well as swimming thrice under it. And we left Portland in the early afternoon, heading for Seattle and a celebration of Adrian Raftery’s career.
JSM 2024, Portland, Day 3
Posted in pictures, Running, Statistics, Travel, University life with tags AI, American Statistical Association, Arianna Rosenbluth, ASA, bandits, Bayesian lasso, Bayesian model averaging, Bayesian model choice, Bayesian nonparametrics, classification, Committee of Presidents of Statistical Societies, completely random measures, convergence diagnostics, COPSS Presidents' Award, data science, David Blackwell, EM algorithm, generalised Bayesian inference, Gibbs measure, Ising model, jISBA, Joint Statistical Meeting, JSM 2024, missing species problem, Monte Carlo EM, Mount Hood National Forest, open water swimming, Oregon, Portland, Rutgers University, spike-and-slab prior, stochastic localization, Thompson sampling, United States of America, Willamette River on August 9, 2024 by xi'an
Bayesian contributed session as the first round of the third day (with a choice of five parallel sessions featuring Bayesian topics!!, actually easier to pick than among the following eight parallel sessions of the 10:30 schedule!!!), with a talk by Tahir Ekin on adversarial outlier detection that could connect with our Oceaner(c) privacy concerns. Then one involving spike & slab (a theme to figure prominently in this special day!!) in mixed response models by Sameer Deshpande, seeking a (unBayesian!) MAP for a latent variable model by Monte Carlo EM. Followed by a talk by Yunyi Shen on completely random measures for estimating the (distribution of the) number of species in heterogeneous populations. Next, Valentin Zulj on (frequentist rather than) Bayesian stacking, on estimating optimal weights for model averaging (which should be posterior probabilities in a pure Bayesian mindframe), including a score function that could lead to generalised Bayesian inference on said weights. Finishing with a talk by Chaegeun Song on correcting Bayesian credible sets towards (frequentist, again!!!) exact coverage for classification (which reminded me of my very first paper with George on correcting frequentist confidence for Binomial observations). With which I could not really engage as seeking a specific coverage level did not seem relevant, imho, but I appreciated the wheel plot representation.
My second morn session was about modern (what else?!) sampling algorithms, although I spent the first dozen minutes wondering whether or not I had entered the wrong room. Until Tianhao Wang focussed on Thompson sampling for bandits. It did prove far enough from my interest for my (sleep deprived) attention to drift too quickly. Only the talk by Yuchen Wu on a spike & slab (as suits the day!) challenge captured enough this wandering attention. Crossing further into my realm of primary topics by considering a target distribution that is a product of distributions. But I did not get from her presentation how a product measure decomposition was inducing higher efficiency (and did not find answers within the arXived preprint). Unless it exploited specific features of the target, like conditional independence between the components. The last talk was by Brice Huang on sampling low temperature Gibbs measures using stochastic localisation.
After coming upon a row of food trucks across the conference centre and being unfairly attracted by an Ethiopian injera picture into a terrible wrap, I returned for the Skeptical about AI session, just a few minutes late, only to find accessing the session was impossible! Quite sad to miss the presentations and the arguments (even though I had heard a previous talk by Genevera Allen when visiting Rutgers two years ago). As a second best, I then joined the recent (of course!) Advances in Bayesian Computation (aka ABC?!) session with a medley of topics, including a data subset versus data sketching model reduction by Sudipto Saha. Which could have consequences on our privacy strategies. And marginal evidence estimation for the Bayesian Lasso by Christopher Hans while avoiding data completion. And another latent variable model with a sequential variational Bayes approach by Bao Anh Vu, using at one point Cappé et al. (2005) EM-based approximation to the log likelihood gradient. Finishing by a back-to-the-future talk by Luke Duttweiler on MCMC convergence diagnostics. Comparing several chains via proximity maps that themselves require some preliminary knowledge about the MCMC kernel. (Nice title though, “the traceplot thickens”!)
The crux of the day was however the 2024 COPSS Award ceremony with several friends featuring among the recipients, Danielle Durante for the Emerging Leaders Award, Regina Liu for the Elizabeth L. Scott Award and Veronika Rockova for the Presidents’ Award. Congrats!!!



JSM 2024, Portland, Day 2
Posted in pictures, Running, Statistics, Travel, University life with tags American Statistical Association, ASA, conference centre, confidence distribution, COPSS Elizabeth L. Scott Award, data science, doubly intractable posterior, evolutionary Monte Carlo, fusion, gene expression, hidden networks, HMC, Joint Statistical Meeting, JSM 2024, Mount Hamilton, Mount Hood National Forest, open water swimming, Oregon, PDMP, Portland, ranger station, SNIP, unification, United States of America, US Department of Agriculture, Willamette River, Zigzag, zigzag algorithm on August 7, 2024 by xi'an
By happenstance, I started my day in the cybersecurity session. With (again) hardly a soul in the room… A first talk on avoiding herding and achieving asymptotic truth learning (about a binary outcome) in a graph by putting constraints on the graph structure, without any clear connection with statistics or cybersecurity. Even less for the second talk on optimising masks. Only with the third one came cybersecurity motivations, the focus being on a two-player Stackelberg game already used in this framework. The result proper was about estimating the parameter of a (rather unrealistic) posited model reproducing the adversarial actions. The last talk about jailbreak attacks was again off-field by miles.

Then attended (as intended!) the 2024 Blackwell-Rosenbluth Award session, featuring the nominees Sharmistha Guha (Texas A&M) on multiple network inference, Simon Mak (Duke) on using Bayesian surrogate models, Guanyang Wang (Rutgers) who recently spoke at our mostly Monte Carlo seminar, Akihiko Nishimura (John Hopkins) on a unification of HMC and PDMPs, most appropriate when located next to Mount Hamilton and the Zigzag river!, and Maria Skoularidou (MIT) on evolutionary Monte Carlo for gene expression. Alas with hardly anyone in the room.

In the afternoon, I went to the (very well-attended this time!, with no seat available for many attendees, incl. yours truly!) COPSS Elizabeth L. Scott Lecture by my friend from Rutgers, Regina Liu, on the highly relevant challenge of combining inferences from diverse data sources. Using the (definitely Rutgerian!) approach of confidence distributions!
Last (late) afternoon, I went swimming from Kevin Duckworth dock, just below the conference centre. Water was quite warm (and green), with a few other swimmers, and no stomachical after-effect, so far. Hence I returned there once again this afternoon.
JSM 2024, Portland, Day 1
Posted in pictures, Running, Statistics, Travel, University life with tags air conditioning, American Statistical Association, ASA, Census Bureau, conference centre, data privacy, data science, differential privacy, disclosure risk, Joint Statistical Meeting, JSM 2024, misinformation, missing-at-random model, Nature, NIST, Oregon, Portland, quantum computers, qubit, rand, Schrödinger's cat, Shannonś information, synthetic data, United States of America, US Department of Agriculture, US Department of Commerce, Willamette River on August 6, 2024 by xi'an
Strolling through the Oregon Conference Centre on 5 Aug, I am as always amazed at how the JSM conference centres have the ability to swallow in thousands of participants without giving an impression of overcrowding! (And appreciating the moderate air conditioning, which for once does not require wearing a fleece indoors!), I must admit that my first impressions of the city itself have been rather poor, as I was (unsuccessfully) seeking an after-hour grocery near the conference centre, I walked through run-down areas and kept passing homeless people, most in a sorry state. And hearing throughout the night, And again this morning in the warehouse maze I jogged through before hitting the Willamette River path, which goes uninterrupted for miles. And showed me an unexpected spot, the Kevin Duckworth dock, where swimming the Willamette is feasible. Hopefully attempting a morn swim before I leave Portland.

I attended the quantum computing session, with a rather light introduction without bringing much light on the nature of qubits for storage and computing. In particular, the role of the complex coefficients of the 0 and 1 states. Then my friend Brani Vidakovic gave us an hand-on demonstration (almost hand-on as the conference facilities were unable to let a code run live!, eons away from quantum performances!!). Using Qiskit and Anaconda. He pointed out that measuring a qubit is destroying its quantum nature, the very equivalent of Schrôdinger’s cat. But does not make it clear whether or not the qubit later returns to a random entity, since frequency stabilisation is assumed, witness the histograms displayed by Brani. or a Nature paper of last year.
Went to a (poorly attended) privacy panel session next. With a defence of the value of differential privacy (Jordan Awan, Purdue) somewhat connected with our own (Ocean) work, although utility not understood in a decision-theoretic sense. And privacy remaining in an one-suits-all sense. With Michael Hawes from the US Census Bureau introducing more dimensions than mere DP (coarsening, suppression, swapping, &tc.), with legal aspects (Title 13) of disclosure risk. With the positive (for me) notion of providing a meaningful assessment of disclosure, with some records being more vulnerable than others. And another one on the cumulative disclosure risk over time. If short in quantitative entries. With Valbona Bejleri (USDA NASS, with a Census of their own) on cell suppressions and metrics for assessing disclosure, eg Shannon’s information entropy. And with Gary Howarth (Privacy Engineering Program, NIST). whose point remained rather unclear to me, like the apparently obvious point that adding features increase dispersion between populations. Although presenting tools for convincing experts and actors of the efficiency of privacy protection techniques.
Which continued (sort of) on the afternoon with a synthetic data for preserving privacy panel session. With Bradley Malin (Vanderbilt U) on the dangers of (Nature Communication paper of 2022). And Harrison Quick using a posterior predictive to achieve differential privacy, albeit considering the posterior predictive as the statistical analysis outcome does not seem the right focus (and recoup my earlier criticism of differential privacy requiring bending one’s prior beliefs). Indeed, as a Bayesian aiming at inference rather than merely at not releasing raw data, I would use synthetic data generated from that posterior predictive to return a posterior on the parameters of interest. And Joshua Snoke (RAND) with cautionary warnings. Like producing “invalid” inference because the reliance on a specific model (but isn’t that the case for most of statistics?). And Roee Gutman (Brown U), discussing the special case of record linkage. More into the difficulties in creating synthetic data. (Like the issue with missingness.) Somehow disappointing in not reaching a more statistical and quantitative perspective, eg by sticking to a Bayesian perspective the whole way.
On a non-technical side, I am surprised at hardly anyone adopting the bring-your-own-container policy in the OCC cafés, or outside, given the large number of attendees carrying one or several liquid containers.