Archive for JMLR

[Split] Frontiers in Statistical Machine Learning [reposted]

Posted in pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , on September 19, 2026 by xi'an

In connection with the IMS conference ICSDS 2026, an IMS Frontiers in Statistical Machine Learning (FSML) satellite workshop takes place on Monday, December 14, 2026 (also) in Split, Croatia, the day before the main conference.

This year’s themes are generative and foundation models for statistics, and the science of deep learning. The keynote speakers are Yuxin Chen, Alexander Henzi, Andrej Risteski, Pragya Sur, Yan Shuo Tan, and Yuexi Wang.

There are two ways to present a poster, both non-archival:

– Workshop Track: short papers of 3 to 5 pages, work in progress welcome. Ten US$500 travel awards for students and postdocs.
– Fast Track: papers already accepted at NeurIPS, ICLR, AISTATS, ICML, UAI, JMLR, or TMLR since August 2025. No additional review.

The deadline for both tracks is Monday, October 19

FSML 2026 organizers are:
Yuansi Chen, ETH Zurich
Sophie Langer, Ruhr University Bochum
Feng Liu, University of Melbourne
Xinwei Shen, University of Washington
Susan Wei, Monash University

comments from Bob

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , on April 10, 2026 by xi'an

Bob replied to my short post with further items of information that I find worth sharing:

Thanks for the kind post, Christian. It’s amusing to be the subject of one of these posts given how many of them I’ve read about other people. I really appreciate your summaries. And thanks to everyone in the audience for all the great feedback during and after the talk. Here’s a link to my slides.

One of your students or postdocs mentioned an approach that does continuous adaptation on some kind of polynomial schedule that is provably correct, but I didn’t manage to write down the author/reference or the name of the person who recommended it. If you happen to know what that is, I’d be grateful for the reference.

I would also like to follow up on the Robert & Andrieu paper you mention, but I could not find the exact reference on your Google Scholar page. The closest match I can find is:

Controlled MCMC for optimal sampling. 2001. C Andrieu, CP Robert. INSEE.

Section 1.3 is titled “Criteria for local adaptation.” The section cites two things. The first is Haario et al.’s (1999) sliding window approach, for which HMC moves too fast to be useful locally. The second is the multiple try approach of Liu et al. (2000) and the delayed rejection approach Tierny and Mira (1999). We applied delayed rejection to HMC step size adaptation in a couple of papers before developing GIST (Modi, Barnett and Carpenter in Bayesian Analysis; Turok, Modi, and Carpenter in AISTATS); these mirror our second GIST paper and third GIST paper in doing the step size adaptation for a whole trajectory and at each leapfrog step. The GIST approach is easier to understand, easier to describe mathematically, easier to implement, and is more efficient.

The nice part about GIST compared to Riemannian HMC is that we do not need to do any volume adjustments (which must be autodiffed through), which are cubic, and we do not need an implicit integrator, which is incredibly fussy to tune. The tradeoff is the we require reversibility of the adaptation, which I think is going to be tricky with varying curvature. Of course, we can’t afford to compute Hessian matrices in high dimensions, but we could manage Hessian-vector products if we could figure out how to use just those and we could also manage low-rank plus diagonal approximations or sketches as described in the Nutpie paper.

We’ve arXived the Nutpie paper since the talk:

Preconditioning HMC by minimizing Fisher divergence. arXiv. 2026. Seyboldt, Carlsen, and Carpenter.

The WALNUTS paper has been accepted by JMLR, but currently only the arXiv version is available:

The within-orbit adaptive leapfrog no-U-turn sampler. 2026. Nawaf Bou-Rabee, Bob Carpenter, Tore Selland Kleppe, Sifan Liu. 2025 arXiv; 2026 to appear JMLR.

Working with Nawaf and Tore has made all the difference in the world on this—it’s not something I could have done by myself. Sifan’s the one who came up with the nice characterization of NUTS and Nawaf’s done a number of additional things like providing mixing time bounds for NUTS (with Milo Marsden, who’s sadly no longer with us—he’s gone into finance).

Furthermore, you can adjust the U-turn criterion from 180 degrees to whatever you want to control how much of a full orbit you get. Those tend to be even more wasteful of iterations, though—this is what the plot from the expected integration time of NUTS is supposed to show, but it was confusing in the talk.

The approach you took with Wu Chengye to randomize number of leapfrog steps made a deep impression on me. It’s also wasteful in leapfrog steps because any number of steps greater than or less than about 1/4 of an orbit is wasteful either in computation or because it leads to more diffusive sampling. You can see that it is roughly as gradient efficient as NUTS in a 1000-dimensional standard normal. Interestingly, it’s worse than NUTS for parameter estimates and better for squared parameter estimates, which is overall a win. Nawaf has also published on randomized HMC. I think we could turn down NUTS U-turn criterion below 180 degrees to get something similar with NUTS, but I haven’t tried it.

One important property of your randomized approach is that it is much much easier to code efficiently for GPUs than NUTS, because the conditionals in NUTS are hard to execute in SIMD fashion. There’s a very nice introduction to this problem by Sountsov, Carroll, and Hoffman, in their paper “Running Markov Chain Monte Carlo on Modern Hardware and Software,” which is out on arXiv and also going into the next edition of the Handbook of MCMC). The thing to read about how to code NUTS on GPU is Dance, Glaser, Orbanz, and Adams’s paper, “Efficiently Vectorized MCMC on Modern Accelerators,” which is on arXiv and ICML 2025.

You can also randomize step size to vary the integration time and avoid harmonics, e.g.,

Randomized Hamiltonian Monte Carlo. 2017. Bou-Rabee and Sanz-Serna. Annals of Applied Probability.

journal-to-conference track at AISTATS 2025!

Posted in Statistics with tags , , , , , , , , , , , , , , on February 18, 2025 by xi'an

An interesting initiative from the organisers of AISTATS 2025, namely, the creation of a journal-to-conference track for papers published in

  • Annals of Statistics
  • Biometrika
  • Journal of the American Statistical Association
  • Journal of the Machine Learning Research
  • Journal of the Royal Statistical Society Series B

after 01 January 2023. They are be considered for a poster presentation at the upcoming AISTATS 2025 conference, held on 03-05 May in Mai Khao, Thailand, as an in-person event.. The submission details are available on the conference webpage. The chairs for this initiative are Pierre Alquier and Kamélia Daudel and the deadline is 15 March.

differential privacy for Bayesian inference

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , on July 30, 2024 by xi'an

As I was reading it in preparation for my JSM²⁴ lecture, I found anew that, in this landmark paper of Dimitrikakis et al. (2017), some limitations of the concept of differential privacy were most apparent:

– a requirement to bend both the model and the prior to fit differential privacy, like switching to Lipschitz constraints or using new (e.g., truncated) priors, which runs contrary to Bayesian principles, although the former can be seen as a form of randomization akin to ABC when the randomization itself is accounted for in the derivation of the “exact” posterior distribution (as in the paper of Berah, Favaro, and Rao (2023) on running MCMC for Bayesian non-parametric estimation on privatized (noisy) data I discussed a few days ago);

– a subtle switch of the randomness from the (privatization) procedure itself (as in Dwork (2006)) to the uncertainty about the parameter, not that it clashes per se with Bayesian principles (even though there is an unclear randomness statement in Theorem 9, when the prior itself seems to become random (?)). Which actually means that producing one realisation from the posterior is the (privatization) procedure, as I realised when discussing with Shenggang Hu in Warwick;

– a linear degradation of the privacy parameter ε when moving from one realisation of the posterior to a simulated sample, assuming iid realisations (I wonder whether or not releasing a dependent sample could involve the ESS instead of the number of MCMC iterations). The upper bound means that no privacy whatsoever is guaranteed for an infinite posterior sample, hence for delivering de facto the posterior (despite the paper producing an (ε,0) bound on the Kullback-Leibler measure of the difference between posteriors);

– an absence of prior knowledge or modelling on the data itself, unless the distance from the actual data to an hypothetical alternative, ρ(x,y), can be interpreted as minus a score function conditional on the actual data, eg the opposite of the log predictive

  • the occurrence of an “exponential prior” that is exp p/m ell (theta) with ell a Lipschitz constant for the associated likelihood

  • the assumption that the data user is not an adversary of the data keeper, in that their utility function is about the parameter (and publicly available)

mostly MC [April]

Posted in Books, Kids, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , on April 5, 2024 by xi'an