Archive for ICML 2023
Bayesian inference and conformal prediction
Posted in Books, Kids, Statistics, University life with tags AISTATS 2021, École Polytechnique, Bayesian inference, conformal prediction, differential privacy, distributed Bayesian inference, federated learning, ICML 2023, large scale inference, Paris-Saclay campus, PhD thesis, thesis defence, uncertainty quantification on October 10, 2023 by xi'andifferentially private distributed Bayesian linear regression with MCMC
Posted in Books, pictures, Statistics, University life with tags Bayesian inference, differential privacy, elephant, ELLIS network, HEC, ICML 2023, Jouy-en-Josas, MCMC, randomisation, unconference on August 30, 2023 by xi'an
An ICML 2023 paper by Barıs¸ Alparslan, Sinan Yıldırım¸ and Ilker Birbil that (re)addresses the issue of privacy when running a Bayesian regression analysis. Resorting to the common notion of differential privacy, imposing a limited variability if a single observation is modified, and a Gaussian randomisation of the observations.
“A differentially private algorithm constrains the difference between the probability distributions of the output values obtained from neighbouring data sets”
In the super classical setup of simple Normal linear regression, y=Xθ+σε. Summary statistics are chosen as
S=X’X and z=X’y,
(why the separation?) then randomised. (Keeping Ŝ definite positive? Not necessarily, it appear.) Inspired directly from Dwork & al. (2014). The authors still manage to spend an entire column in (re)deriving the conditional Normal distribution of z conditional on S and (θ,σ)… Which is later exploited for integrating z out in the MCMC algorithm.
“some important differences between our work and that of Bernstein & Sheldon (2019) [stem] from the choice of summary statistics and the consequent hierarchical structure used for modelling linear regression [and]lead to significant differences in the inference methods as well as significant computational advantages [O(d³) vs. O(d⁶)]”
In a distributed setting several agents are handling their own data and keep their privacy by the same mechanishttps://www.slideshare.net/xianblog/discussion-of-icml23pdfm [as in the top graph from the paper]. On principle, a Bayesian analysis of the resulting hierarchical model should directly consider the posterior on the global parameter by considering the distributions of the randomised pairs (ẑ,Ŝ). The elephant in the room is the distribution of the regressors, which is customarily unknown and not accounted for in a traditional Bayesian analysis. It is needed here due to the division in S and z, plus the randomisation step that calls for the posterior distribution of S given Ŝ. Elephant that is exfiltrated by either assuming Normality or substituting Ŝ for S without accounting for the noise! Definitely not exactly Bayesian. Another column is spent on the Metropolis-within-Gibbs simulation of the posterior…
Overall, I remain reserved about this approach, since it does not follow a clear Bayesian pathway and in particular does not incorporate privacy as part of the Bayesian decision analysis.
exact yet private MCMC
Posted in Statistics with tags Arrowleaf Cellars, differential privacy, ergodicity, ICML 2023, Lake Okanagan, MCMC, Metropolis-Hastings algorithm, Okanagan vineyards, Poisson subsampling, reversibility, spectral gap, stationarity on August 9, 2023 by xi'an
“at each iteration, DP-fast MH first samples a minibatch size and checks if it uses a minibatch of data or full-batch data. Then it checks whether to require additional Gaussian noise. If so, it will instantiate the Gaussian mechanism which adds Gaussian noise to the energy difference function. Finally, it chooses accept or reject θ′ based on the noisy acceptance probability.”
Private, Fast, and Accurate Metropolis-Hastings for Large-Scale Bayesian Inference is an(other) ICML²³ paper, written by Wanrong Zhang and Ruqi Zhang. Who are running MCMC under DP constraints. For one thing, they compute the MH acceptance probability with a minibatch, which is Poisson sampled (in order to guarantee privacy). It appears as a highly calibrated algorithm (see, e.g., Algorithm 1). Under the assumption (1) that the difference between individual log densities for two values of the parameter is upper bounded (in the data), differential privacy is established as failing to detect for certain a datapoint from the MCMC output. Interestingly, the usual randomisation leading to pricacy is operated on the energy level, rather than on observations or summary statistics, although this may prove superfluous when there is enough randomness provided by the MH step itself: “inherent privacy guarantees in the MH algorithm”
“when either the privacy hyperparameter ϵ or δ becomes small, the convergence rate becomes small, characterizing how much the privacy constraint slows down the convergence speed of the Markov chain”
The major results of the paper are privacy guarantees (at each iteration) and preservation of the proper target distribution, in contrast with earlier versions. In particular, adding the Gaussian noise to the energy does not impact reversibility. (Even though I am not 100% sure I buy the entire argument about reversibility (in Appendix C) as it sounds too easy!) The authors even achieve a bound on the relative spectral gaps.
ellis unconference [not in Hawai’i]
Posted in pictures, Running, Travel, University life with tags Bièvre, business school, Chateaubriand, CIRM, diffusions, ELLIS network, Europe, Flatiron Institute, France, Hawaii, HEC, Hi! Paris, ICML 2023, International Conference on Machine Learning, ISBA 2021, Jouy-en-Josas, Maurice Kenneth Tweedie, mirror workshop, normalising flow, Paris, Paris Artificial Intelligence for Society, Paris Artificial Intelligence Research Institute, SMC, the European Laboratory for Learning and Intelligent Systems, Tweedie's formula, unconference, variational Bayes methods, Verrières, warping, Wasserstein distance on July 26, 2023 by xi'an
As ICML 2023 is happening this week, in Hawai’i, many did not have the opportunity to get there, for whatever reason, and hence the ellis (European Lab for Learning {and} Intelligent Systems] board launched [fairly late!] with the help of Hi! Paris an unconference (i.e., a mirror) that is taking place in HEC, Jouy-en-Josas, SW of Paris, for AI researchers presenting works (theirs or others’) presented at ICML 2023. Or not. There was no direct broadcasting of talks as we had (had) in CIRM for ISBA 2020 2021. But some presentations based on preregistered talks. Over 50 people showed up in Jouy.
As it happened, I had quite an exciting bike ride to the HEC campus from home, under a steady rain, crossing a (modest) forest (de Verrières) I had never visited before, despite it being a few km from home, getting a wee bit lost, stopped by a train Xing between Bièvre and Jouy, and ending up at the campus just in time for the first talk (as I had not accounted for the huge altitude differential). Among curiosities met on the way, “giant” sequoias, a Tonkin pond, Chateaubriand’s house.
As always I am rather impressed by the efficiency of AI-ML conferences run, with papers+slides+reviews online, plus extra material as in this example. Lots of papers on diffusion models this year, apparently. (In conjunction with the trend observed at the Flatiron workshop last Fall.) Below are incoherent tidbits from the presentations I attended:
- exponential convergence of the Sinkhorn algorithm by Alain Durmus and co-authors, with the surprise occurrence of a left Haar measure
- a paper (by Jerome Baum, Heishiro Kanagawa, and my friend Arthur Gretton) on Stein discrepancy, with an Zanella Stein operator relating to Metropolis-Hastings/Barker since it has expectation zero under stationarity, interesting approach to variable length random variables, not a RJMCMC, but nearby.
- the occurance of a criticism of the EU GDPR that did not feel appropriate for synthetic data used in privacy protection.
- the alternative Sliced Wasserstein distance, making me wonder if we could optimally go from measure μ to measure ζ using random directions or how much was lost this way.

- Information Maximizing Optimal Transport with dubious substitute for conditional expectation:
as (a) densities are replaced with kernel estimates, (b) the outer density may be very small, (c) no variance assessment is provided.

- Markov score climbing and transport score climbing using a normalising flow, for variational approximation, presented by Christian Naesseth, with a warping transform that sounded like inverting the flow (?)
- Yazid Janati not presenting their ICML paper State and parameter learning with PARIS particle Gibbs written with Gabriel Cardoso, Sylvain Le Corff, Eric Moulines and Jimmy Olsson, but another work with a diffusion based model to be learned by SMC and a clever call to Tweedie’s formula. (Maurice Kenneth Tweedie, not Richard Tweedie!) Which I just realised I have used many times when working on Bayesian shrinkage estimators
[A]ABC in Hawai’i
Posted in Statistics with tags 5th Symposium on Advances in Approximate Bayesian Inference, AABI, ABC, ABC in, approximate Bayesian inference, Hawaii, ICML 2023, Vietnam, workshop on April 6, 2023 by xi'an

“at each iteration, DP-fast MH first samples a minibatch size and checks if it uses a minibatch of data or full-batch data. Then it checks whether to require additional Gaussian noise. If so, it will instantiate the Gaussian mechanism which adds Gaussian noise to the energy difference function. Finally, it chooses accept or reject θ′ based on the noisy acceptance probability.”