Sirio Legramanti, Daniele Durante, and Pierre Alquier just arXived a massive paper on the concentration of discrepancy–based ABC posteriors via Rademacher complexity, which includes MMD and Wasserstein distance-based ABC methods. The paper provides sufficient conditions under which a discrepancy within the integral probability semimetrics class guarantees uniform convergence and concentration of the induced ABC posterior, without necessarily requiring suitable regularity conditions for the underlying data generating process and the assumed statistical model, meaning that they also cover misspecified cases. In particular, the authors derive upper and lower bounds on the limiting acceptance probabilities for the ABC posterior to remain well–defined for a sample size large enough. They thus deliver an improved understanding of the factors that govern the uniform convergence and concentration properties of discrepancy–based ABC posteriors under a fairly unified perspective, which I deem a significant advance on the several papers my coauthors Ernst Bernton, David Frazier, Mathieu Gerber, Pierre Jacob, Gael Martin, Judith Rousseau, Robin Ryder, and yours truly produced in that domain over the past years (although our Series B misspecification paper does not appear in the reference list!)
“…as highlighted by the authors, these [convergence] conditions (i) can be difficult to verify for several discrepancies, (ii) do not allow to assess whether some of these discrepancies can achieve convergence and concentration uniformly over P(Y), and (iii) often yield bounds which hinder an in–depth understanding of the factors regulating these limiting properties”
The first result is that, asymptotically in n and a fixed large-enough tolerance, the ABC posterior is always well–defined but within a Rademacher ball of the pseudo-true posterior, larger than the tolerance ε when the Rademacher complexity does not vanish in n (a feature on which my intuition is found to be lacking!, since it seems to relate solely to the class of functions adopted for the definition of said discrepancy). When the tolerance ε(n) decreases to its minimum, as in our paper, the speed of concentration is similar to ours, with a speed slower than √n. And assuming the tolerance ε(n) decreases to its minimum slower than √n but faster than the Rademacher complexity.
“…the bound we derive crucially depends on [the Rademacher complexity], which is specific to each discrepancy D and plays a fundamental role in controlling the rate of concentration of the ABC posterior.”
The paper also opens towards non-iid settings (as in our Wasserstein paper) and generalized likelihood–free Bayesian inference à la Bissiri et al. (2016). A most interesting take on the universality of ABC convergence, thus, although assuming bounded function spaces from the start.
The main BayesComp meeting started right after the ABC workshop and went on at a grueling pace, and offered a constant conundrum as to which of the four sessions to attend, the more when trying to enjoy some outdoor activity during the lunch breaks. My overall feeling is that it went on too fast, too quickly! Here are some quick and haphazard notes from some of the talks I attended, as for instance the practical parallelisation of an SMC algorithm by Adrien Corenflos, the advances made by Giacommo Zanella on using Bayesian asymptotics to assess robustness of Gibbs samplers to the dimension of the data (although with no assessment of the ensuing time requirements), a nice session on simulated annealing, from black holes to Alps (if the wrong mountain chain for Levi), and the central role of contrastive learning à la Geyer (1994) in the GAN talks of Veronika Rockova and Éric Moulines. Victor Elvira delivered an enthusiastic talk on our massively recycled importance on-going project that we need to complete asap!
In a not-solely-ABC session, I appreciated Sirio Legramanti speaking on comparing different distance measures via Rademacher complexity, highlighting that some distances are not robust, incl. for instance some (all?) Wasserstein distances that are not defined for heavy tailed distributions like the Cauchy distribution. And using the mean as a summary statistic in such heavy tail settings comes as an issue, since the distance between simulated and observed means does not decrease in variance with the sample size, with the practical difficulty that the problem is hard to detect on real (misspecified) data since the true distribution behing (if any) is unknown. Would that imply that only intrinsic distances like maximum mean discrepancy or Kolmogorov-Smirnov are the only reasonable choices in misspecified settings?! While, in the ABC session, Jeremiah went back to this role of distances for generalised Bayesian inference, replacing likelihood by scoring rule, and requirement for Monte Carlo approximation (but is approximating an approximation that a terrible thing?!). I also discussed briefly with Alejandra Avalos on her use of pseudo-likelihoods in Ising models, which, while not the original model, is nonetheless a model and therefore to taken as such rather than as approximation.