After a final early morning run in the rising sun, reaching a point north of the city where I could see a nearby volcano (not Fuji-san!), I attended both morning on MCMC, with (again) a range of interesting, mostly novel, questions and solutions. With no overlap with his talk at mostly Monte Carlo last month, Sam Livingstone gave convincing motivations for using new tools for designing proper scaling (in the limit) for some adaptive MCMC. Charles Margossian’s talk was particularly exciting for pushing for single step MCMC when massively run in parallel, provided warm-up is over (enough). And Saif Syed discussing a new version of annealed SMC, using the variability of the estimated normalising constant as an assessment of the (although I could not catch how the tridimensional calibration was handled). I was less convinced by Kyle Kuang’s approach to overcome identifiability issues such as label switching, as it sounded too simple to be universally applicable. Bingjing Tang returned to the challenge of doubly intractable posteriors. With motivations from functional inference and a solution reminding me of noise contrastive estimation. until I spoke with the authors and realised it was much closer to our recent paper with Edoardo and Julien. While Bjorn Sprungk’s talk on Metropolized interacting particle sampling reminded me of our pinball sampler, presented… 30 years ago at the 1996 Valencia meeting! But using the product of posteriors as a target sounds suboptimal when a target that would keep particles apart (with the correct marginals) would prove more exploratory.
The noon break was the last opportunity to sample one of the food stalls in the fish market nearby, with a spicy curry udon bowl. Cutting the eel addiction!
My final session—before catching a shinkansen to Tokyo for the Information Geometry, Privacy and Monte Carlo ISBA Satellite Meeting at the Institute of Statistical Mathematics—was about loss-based posteriors, with our PhD student Shreya Roy presenting her work on prequential posteriors. And Kshitij Khare on using a loss that allows for a regular Gibbs sampler implementation via a pseudo-model and consistency properties.
This cuvée of ISBA World Meeting was exceptionally (gouleyante and) enjoyable (except for my recurrent sleeping issues) from the diverse and well-balanced programme, to the choice of plenary speakers, to the practicality of the conference centre (except for the queues for the lift!) and its location in Nagoya, with its own, unsuspected, perks! With no food poisoning this time!! ISBA 2028 is scheduled to take place in Milwaukee and I am very unlikely to attend, unless a rogue mirror pops up!
The main BayesComp meeting started right after the ABC workshop and went on at a grueling pace, and offered a constant conundrum as to which of the four sessions to attend, the more when trying to enjoy some outdoor activity during the lunch breaks. My overall feeling is that it went on too fast, too quickly! Here are some quick and haphazard notes from some of the talks I attended, as for instance the practical parallelisation of an SMC algorithm by Adrien Corenflos, the advances made by Giacommo Zanella on using Bayesian asymptotics to assess robustness of Gibbs samplers to the dimension of the data (although with no assessment of the ensuing time requirements), a nice session on simulated annealing, from black holes to Alps (if the wrong mountain chain for Levi), and the central role of contrastive learning à la Geyer (1994) in the GAN talks of Veronika Rockova and Éric Moulines. Victor Elvira delivered an enthusiastic talk on our massively recycled importance on-going project that we need to complete asap!
In a not-solely-ABC session, I appreciated Sirio Legramanti speaking on comparing different distance measures via Rademacher complexity, highlighting that some distances are not robust, incl. for instance some (all?) Wasserstein distances that are not defined for heavy tailed distributions like the Cauchy distribution. And using the mean as a summary statistic in such heavy tail settings comes as an issue, since the distance between simulated and observed means does not decrease in variance with the sample size, with the practical difficulty that the problem is hard to detect on real (misspecified) data since the true distribution behing (if any) is unknown. Would that imply that only intrinsic distances like maximum mean discrepancy or Kolmogorov-Smirnov are the only reasonable choices in misspecified settings?! While, in the ABC session, Jeremiah went back to this role of distances for generalised Bayesian inference, replacing likelihood by scoring rule, and requirement for Monte Carlo approximation (but is approximating an approximation that a terrible thing?!). I also discussed briefly with Alejandra Avalos on her use of pseudo-likelihoods in Ising models, which, while not the original model, is nonetheless a model and therefore to taken as such rather than as approximation.