Archive for International Conference on Machine Learning
persuasive privacy at ICML 2026
Posted in Books, Statistics, Travel, University life with tags #ERCSyG, Bayesian decision theory, differential privacy, game theory, ICML 2026, International Conference on Machine Learning, Korea, Ocean, persuasive privacy, poster, privacy, probabilistic differential privacy, Seoul, Stackelberg game, Université Paris Dauphine on July 7, 2026 by xi'anBayesian persuasive privacy at ICML²⁶
Posted in Books, Mountains, pictures, Statistics, Travel, University life with tags #ERCSyG, adversarial privacy, Bayesian inference, Bayesian privacy, ICML 2026, International Conference on Machine Learning, LLM, LLM reviewing, Ocean, Oceanerc, peer review process, persuasive privacy, proceedings, Seoul, South Korea, statistical machine learning, watermarking on May 14, 2026 by xi'anellis unconference [not in Hawai’i]
Posted in pictures, Running, Travel, University life with tags Bièvre, business school, Chateaubriand, CIRM, diffusions, ELLIS network, Europe, Flatiron Institute, France, Hawaii, HEC, Hi! Paris, ICML 2023, International Conference on Machine Learning, ISBA 2021, Jouy-en-Josas, Maurice Kenneth Tweedie, mirror workshop, normalising flow, Paris, Paris Artificial Intelligence for Society, Paris Artificial Intelligence Research Institute, SMC, the European Laboratory for Learning and Intelligent Systems, Tweedie's formula, unconference, variational Bayes methods, Verrières, warping, Wasserstein distance on July 26, 2023 by xi'an
As ICML 2023 is happening this week, in Hawai’i, many did not have the opportunity to get there, for whatever reason, and hence the ellis (European Lab for Learning {and} Intelligent Systems] board launched [fairly late!] with the help of Hi! Paris an unconference (i.e., a mirror) that is taking place in HEC, Jouy-en-Josas, SW of Paris, for AI researchers presenting works (theirs or others’) presented at ICML 2023. Or not. There was no direct broadcasting of talks as we had (had) in CIRM for ISBA 2020 2021. But some presentations based on preregistered talks. Over 50 people showed up in Jouy.
As it happened, I had quite an exciting bike ride to the HEC campus from home, under a steady rain, crossing a (modest) forest (de Verrières) I had never visited before, despite it being a few km from home, getting a wee bit lost, stopped by a train Xing between Bièvre and Jouy, and ending up at the campus just in time for the first talk (as I had not accounted for the huge altitude differential). Among curiosities met on the way, “giant” sequoias, a Tonkin pond, Chateaubriand’s house.
As always I am rather impressed by the efficiency of AI-ML conferences run, with papers+slides+reviews online, plus extra material as in this example. Lots of papers on diffusion models this year, apparently. (In conjunction with the trend observed at the Flatiron workshop last Fall.) Below are incoherent tidbits from the presentations I attended:
- exponential convergence of the Sinkhorn algorithm by Alain Durmus and co-authors, with the surprise occurrence of a left Haar measure
- a paper (by Jerome Baum, Heishiro Kanagawa, and my friend Arthur Gretton) on Stein discrepancy, with an Zanella Stein operator relating to Metropolis-Hastings/Barker since it has expectation zero under stationarity, interesting approach to variable length random variables, not a RJMCMC, but nearby.
- the occurance of a criticism of the EU GDPR that did not feel appropriate for synthetic data used in privacy protection.
- the alternative Sliced Wasserstein distance, making me wonder if we could optimally go from measure μ to measure ζ using random directions or how much was lost this way.

- Information Maximizing Optimal Transport with dubious substitute for conditional expectation:
as (a) densities are replaced with kernel estimates, (b) the outer density may be very small, (c) no variance assessment is provided.

- Markov score climbing and transport score climbing using a normalising flow, for variational approximation, presented by Christian Naesseth, with a warping transform that sounded like inverting the flow (?)
- Yazid Janati not presenting their ICML paper State and parameter learning with PARIS particle Gibbs written with Gabriel Cardoso, Sylvain Le Corff, Eric Moulines and Jimmy Olsson, but another work with a diffusion based model to be learned by SMC and a clever call to Tweedie’s formula. (Maurice Kenneth Tweedie, not Richard Tweedie!) Which I just realised I have used many times when working on Bayesian shrinkage estimators
scalable Metropolis-Hastings, nested Monte Carlo, and normalising flows
Posted in Books, pictures, Statistics, University life with tags Bayesian neural networks, Bernstein-von Mises theorem, CIF, computing cost, conferences, density approximation, dissertation, doubly intractable posterior, evidence, ICML 2019, ICML 2020, image analysis, International Conference on Machine Learning, L¹ convergence, logistic regression, nesting Monte Carlo, normalising flow, PhD, probabilistic programming, quarantine, SAME algorithm, scalable MCMC, thesis defence, University of Oxford, variational autoencoders, viva on June 16, 2020 by xi'an
Over a sunny if quarantined Sunday, I started reading the PhD dissertation of Rob Cornish, Oxford University, as I am the external member of his viva committee. Ending up in a highly pleasant afternoon discussing this thesis over a (remote) viva yesterday. (If bemoaning a lost opportunity to visit Oxford!) The introduction to the viva was most helpful and set the results within the different time and geographical zones of the Ph.D since Rob had to switch from one group of advisors in Engineering to another group in Statistics. Plus an encompassing prospective discussion, expressing pessimism at exact MCMC for complex models and looking forward further advances in probabilistic programming.
Made of three papers, the thesis includes this ICML 2019 [remember the era when there were conferences?!] paper on scalable Metropolis-Hastings, by Rob Cornish, Paul Vanetti, Alexandre Bouchard-Côté, Georges Deligiannidis, and Arnaud Doucet, which I commented last year. Which achieves a remarkable and paradoxical O(1/√n) cost per iteration, provided (global) lower bounds are found on the (local) Metropolis-Hastings acceptance probabilities since they allow for Poisson thinning à la Devroye (1986) and second order Taylor expansions constructed for all components of the target, with the third order derivatives providing bounds. However, the variability of the acceptance probability gets higher, which induces a longer but still manageable if the concentration of the posterior is in tune with the Bernstein von Mises asymptotics. I had not paid enough attention in my first read at the strong theoretical justification for the method, relying on the convergence of MAP estimates in well- and (some) mis-specified settings. Now, I would have liked to see the paper dealing with a more complex problem that logistic regression.
The second paper in the thesis is an ICML 2018 proceeding by Tom Rainforth, Robert Cornish, Hongseok Yang, Andrew Warrington, and Frank Wood, which considers Monte Carlo problems involving several nested expectations in a non-linear manner, meaning that (a) several levels of Monte Carlo approximations are required, with associated asymptotics, and (b) the resulting overall estimator is biased. This includes common doubly intractable posteriors, obviously, as well as (Bayesian) design and control problems. [And it has nothing to do with nested sampling.] The resolution chosen by the authors is strictly plug-in, in that they replace each level in the nesting with a Monte Carlo substitute and do not attempt to reduce the bias. Which means a wide range of solutions (other than the plug-in one) could have been investigated, including bootstrap maybe. For instance, Bayesian design is presented as an application of the approach, but since it relies on the log-evidence, there exist several versions for estimating (unbiasedly) this log-evidence. Similarly, the Forsythe-von Neumann technique applies to arbitrary transforms of a primary integral. The central discussion dwells on the optimal choice of the volume of simulations at each level, optimal in terms of asymptotic MSE. Or rather asymptotic bound on the MSE. The interesting result being that the outer expectation requires the square of the number of simulations for the other expectations. Which all need converge to infinity. A trick in finding an estimator for a polynomial transform reminded me of the SAME algorithm in that it duplicated the simulations as many times as the highest power of the polynomial. (The ‘Og briefly reported on this paper… four years ago.)
The third and last part of the thesis is a proposal [to appear in ICML 20] on relaxing bijectivity constraints in normalising flows with continuously index flows. (Or CIF. As Rob made a joke about this cleaning brand, let me add (?) to that joke by mentioning that looking at CIF and bijections is less dangerous in a Trump cum COVID era at CIF and injections!) With Anthony Caterini, George Deligiannidis and Arnaud Doucet as co-authors. I am much less familiar with this area and hence a wee bit puzzled at the purpose of removing what I understand to be an appealing side of normalising flows, namely to produce a manageable representation of density functions as a combination of bijective and differentiable functions of a baseline random vector, like a standard Normal vector. The argument made in the paper is that imposing this representation of the density imposes a constraint on the topology of its support since said support is homeomorphic to the support of the baseline random vector. While the supporting theoretical argument is a mathematical theorem that shows the Lipschitz bound on the transform should be infinity in the case the supports are topologically different, these arguments may be overly theoretical when faced with the practical implications of the replacement strategy. I somewhat miss its overall strength given that the whole point seems to be in approximating a density function, based on a finite sample.
