
Speaker: Larry Wasserman (Carnegie Mellon University)

Speaker: Larry Wasserman (Carnegie Mellon University)

On 28 October, I spent the day at Institut d’Astrophysique de Paris (where I used to work on PMC for cosmology between 2005 and 2009), as a committee member for the habilitation defence of Florent Leclercq. Not only it was nice to be back in this unique institution (with vestiges from Laplace’s era), but this was a fantastic habilitation, with a superb thesis that beautifully gathered the different fields mastered by the candidate in a highly coherent discourse. And could serve as an introduction to cosmostatistics for many.

And provided the background to ten years (post-PhD) of research on forward modelling in cosmology and resulting Bayesian statistical analysis either by implicit likelihood (or likelihood-free) inference or by field-level inference. He describes the Simbelmynë software he developed to produce maps of the density field and analyse dark matter dynamics. Ẁhose name is borrowed from Tolkien (along with a quote from Guy Gavriel Kay!):
“How fair are the bright eyes in the grass! Evermind they are called, simbelmynë in this land of Men, for they blossom in all the seasons of the year, and grow where dead men rest.” — J.R.R. Tolkien, The Lord of the Ring
And the Bayesian computational and modelling tools he elaborated, like SELFI (Simulator expansion for likelihood-free inference, Leclercq et al., 2019), that relates to Michael Gutmann’s and Juka Corander’s BOLFI. (Obviously, I did not get every aspect right from just reading the thesis and attending the lecture, in particular the remarks on using SELFI to assess model misspecification. But I remain impressed by the scope of the work and its likely impact on the field!)
Martin Metodiev and his coauthor(es)s have produced another paper on the THAMES Monte Carlo method when specifically targetting marginal likelihoods for mixture models. Since this problem has long been a central interest of mine’s and since the method is closely connected with the harmonic mean solution we developed with Darren Wraith in 2009, (and also included in our 2009 survey with Jean-Michel Marin of evidence approximations, published in Frontiers of Statistical Decision Making and Bayesian Analysis for Jim Berger’s 60th birthday), I quickly went into the paper. The core purpose of this paper is to adapt THAMES to a multimodal setting since using an ellipsoidal region as the support of the Uniform reciprocal importance sampling distribution does not make sense for a multimodal target. After reading it a few times, and while some computational aspects remain obscure to me, I am not convinced this brings an adequate answer to the challenge. Indeed, while the approach borrows directly from Berkhof et al. (2003) that inspired the resolution we proposed, Jeong (Kate) Lee and myself, the issues I have with the current proposal are that
1. the evacuation of earlier methods as not simple or not universal enough is rather disingenuous. For instance, software that do not return (latent) allocation vectors can easily be post-processed. And the current method uses allocation probabilities just the same (in Section 3.3). Similarly, the random shuffling answer to label (lack of) switching proposed by Sylvia Früwirth-Schnatter—which again can be achieved by post-processing—cannot be rejected on the sole basis that the component means (based on the MCMC sample) are all similar. It is furthermore debatable that the current proposal is simple, when involving relabelling à la Stephens, averaging over permutations, selecting over said permutations by constructing a graph over components (section 3.2.1) and running a quadratic discriminant analysis (section 3.2.2) on the posterior sample, based on an arbitrary Normal representation of the distributions of the clusters, and finally defining a new ordering constraint (section 3.2.3). Computing efforts required by the respective methods do not appear in the main text.
2. the handling of the label switching issue—the reason why Larry Wasserman saw mixtures at the same magnitude of evil as tequila!—is problematic for several reasons. My position (since at least 2000!) on the matter is that the proper posterior sample must exhibit label switching and come close to symmetry among the “components”. The label switching problem (section 3.1) is rather when the MCMC sample does not “switch their labels”. The relabelling approach (e.g., à la Stephens) allows for a differentiation between components, to some extent, which helps with computing basic posterior moments for point estimation or for the calibration of the support of the Uniform reciprocal importance sampling distribution, but the use of any relabelling procedure is tampering with the original MCMC sample and thus bound to impact the distribution of the resulting relabelled sample. Furthermore, relabelling depends on the value of G, whereas the actual number of (significant) modes in the posterior is also connected with the (partial) fit of the data to the model, meaning the creation of further modes than those linked with relabelling. Especially when the model is misspecified. Incidentally, the symmetrised version of THAMES (5) does not require relabelling. Neither does the Bayes factor. In addition, the experiment section (4.1.2) mentions that bridge sampling is biased by a factor of G!, which comes as a surprise to me since I associated this factor with the call to Sid Chib’s formula in the absence of label switching, i.e. when the MCMC sample was stuck on a mode, as exposed by Radford Neal in 1999. Is it because bridge sampling is applied to the relabelled sample? It is also surprising that the gap appears in the simulated datasets (Fig.3) and not in the real ones (Fig.5).
3. the (legitimate) purpose of using marginal likelihoods for selecting the number G of components is weakened by the intrusion of alternate proposals to assess G from the data, like the criterion of overlap (section 3.2.1), which instead aims at the number of clusters, with an elimination of “empty components” that should either remain a possibility (within a regular mixture model) or be evacuated with a different modelling (à la Diebolt & Robert, or à la Wasserman). This overlapping criterion is further used in the discriminant analysis that only applies to “non-overlapping components” of the mixture (section 3.2.3)—at which point I got lost in the reordering and simplification of the computation of THAMES (but got reminded of the results of Agostino Nobile in the 2000’s, with whom I used to discuss a lot in my yearly visit to the University of Glasgow).
4. several mentions are made of the other estimators being biased, which is indeed the case for bridge sampling (if not necessarily for importance sampling), but not necessarily a central issue, while the original generalised harmonic proposal by Gelfand and Dey (1994) and thus THAMES produce an unbiased estimator of the inverse of the evidence (thus neither of the evidence nor of the log-evidence). However, in the paper, the volume of the support of the Uniform reciprocal importance sampling distribution is estimated by a basic Monte Carlo coverage probability in (3), which induces the same type of bias as the other methods.
30 May was a day of first and last times, if not in capital ways (for me), on the One World ABC webinar. This was the first time we had a talk by Mark Beaumont and also the first time I had team experience of facing a smoking participant, while this was the last time of our monthly webinar for the (Northern) academic year.
The talk was about model misspecification in population genomic, from an ABC perspective with the motivation of common noticeable difference between the distributions of d(s,s⁰) and d(s,s’), distances between the prior predictively simulated summary statistics and the observed ones vs posterior generated ones, which should indicates misspecification, esp with complicated models. Mark and his coauthors then supported a gradual elimination of summary statistics to diminish the discrepancy, hence voluntarily impoverishing the model. While blaming the statistics sounded a bit like shooting the messenger, the resolution is of obvious interest if backing from modelling the misspecification itself.
Some of the presented work was conducted in Ward et al (2022, NeurIPS) with a reference to the outlying Ratmann et al (2009) we later discussed, for including tolerance as an extra parameter ε, thereby reconsidering Wikinson’s exact ABC for noisy observations y by the medium of a normalising flow on the marginal distribution of the denoised x (learned from the prior predictive)
More precisely, the idea is to drop summary statistics by checking whether or not the observed S⁰ belongs to HPD region, removing one component of S at a time, using e.g. a k-NN estimate for the summary density (hence depending on parameterisation of said statistics for the distance). Hopefully, the process stops before loosing identifiability by using too few statistics. I also wondered at multiple uses of the data in this sequential procedure but Mark argued for adopting a meta- or pragma- or Gelmanian- Bayesian perspective in the end!
Another perk was the appearance of (and illustration with) the Scottish Wildcat, mentioned in The Guardian a few months ago and discussed in the ‘Og, with further papers exploring more aspects of this hybridization, like a posterior applied to a much more complex phylogenic tree reconstruction for cats of different creeds and many related parameters.
Sirio Legramanti, Daniele Durante, and Pierre Alquier just arXived a massive paper on the concentration of discrepancy–based ABC posteriors via Rademacher complexity, which includes MMD and Wasserstein distance-based ABC methods. The paper provides sufficient conditions under which a discrepancy within the integral probability semimetrics class guarantees uniform convergence and concentration of the induced ABC posterior, without necessarily requiring suitable regularity conditions for the underlying data generating process and the assumed statistical model, meaning that they also cover misspecified cases. In particular, the authors derive upper and lower bounds on the limiting acceptance probabilities for the ABC posterior to remain well–defined for a sample size large enough. They thus deliver an improved understanding of the factors that govern the uniform convergence and concentration properties of discrepancy–based ABC posteriors under a fairly unified perspective, which I deem a significant advance on the several papers my coauthors Ernst Bernton, David Frazier, Mathieu Gerber, Pierre Jacob, Gael Martin, Judith Rousseau, Robin Ryder, and yours truly produced in that domain over the past years (although our Series B misspecification paper does not appear in the reference list!)
“…as highlighted by the authors, these [convergence] conditions (i) can be difficult to verify for several discrepancies, (ii) do not allow to assess whether some of these discrepancies can achieve convergence and concentration uniformly over P(Y), and (iii) often yield bounds which hinder an in–depth understanding of the factors regulating these limiting properties”
The first result is that, asymptotically in n and a fixed large-enough tolerance, the ABC posterior is always well–defined but within a Rademacher ball of the pseudo-true posterior, larger than the tolerance ε when the Rademacher complexity does not vanish in n (a feature on which my intuition is found to be lacking!, since it seems to relate solely to the class of functions adopted for the definition of said discrepancy). When the tolerance ε(n) decreases to its minimum, as in our paper, the speed of concentration is similar to ours, with a speed slower than √n. And assuming the tolerance ε(n) decreases to its minimum slower than √n but faster than the Rademacher complexity.
“…the bound we derive crucially depends on [the Rademacher complexity], which is specific to each discrepancy D and plays a fundamental role in controlling the rate of concentration of the ABC posterior.”
The paper also opens towards non-iid settings (as in our Wasserstein paper) and generalized likelihood–free Bayesian inference à la Bissiri et al. (2016). A most interesting take on the universality of ABC convergence, thus, although assuming bounded function spaces from the start.
| David on Bayesian workflow [book review… | |
| Ewan on Tidy up, Claude! | |
| xi'an on Tidy up, Claude! | |
| Ewan on Tidy up, Claude! | |
| Ewan on Já við viðræðum. Já við Evrópu… |
