Archive for tequila

robust simulation-based inference

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , on March 7, 2026 by xi'an

This new arXival by Lorenzo Tomaselli, Valérie Ventura, and Larry Wasserman (from CMU) considers simulation-based inference under model misspecification (as we did for ABC in our 2020 Series B paper). Which is almost always the case. In the paper, SBI is defined as producing N parameters and N samples from the prior and the corresponding sampling distribution, respectively, and then doubling the resulting samples by permuting at random the parameters θ. This means that the second half is distributed from the product of the prior and of the marginal, hence that the classification odds ratio is equal to the likelihood, hence providing an estimation method (andlikelihood trick) à la Geyer. From this estimate, an ABC p-value can be derived, but it is incorrect as such when the model is misspecified. Hence the use of the Hellinger discrepancy, the power divergence and the kernel distance (or MMD) as alternatives to the misspecified MLE.

The paper then expands on approximating density ratios by virtue of a reproducing kernel Hilbert space, using a Gaussian kernel. (With a nice remark on requiring only one single ratio estimator for all values of θ, albeit in the joint space.) And focus on a studentized MMD estimator (à la e-value) to build a confidence set that remains valid under model misspecification. And without regularity assumptions.

Another approach is further explored, based on exponential tilting—of which I am not a great fan, from being highly dependent on the choice of the pseudo-sufficient statistic to require an intractable normalising constant, to requiring an extra optimization, even though I appreciate the mathematical appeal of the construct. Which seems to require a sample simulation for each value of θ at the learning stage, albeit relying on the same likelihood trick. The appropriateness of the tilting can be tested by a goodness of fit test tailored for the SBI structure, which sounds rather greedy in the required simulations. 

Besides the g-and-k distribution example (which, as pointed out several times on the ‘Og, is not intractable, strictly speaking!), the paper studies a mixture example, despite Larry dubbing them as evil as tequila a long while ago! (The paper also offers a section called accoutrements, which is my first encounter with this use of the term, usually found in medieval contexts!)

Note that Larry will present the paper at the OWABI webinar next 25 March!

easily computed marginal likelihoods for multivariate mixture models using the THAMES estimator

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , on May 25, 2025 by xi'an

Martin Metodiev and his coauthor(es)s have produced another paper on the THAMES Monte Carlo method when specifically targetting marginal likelihoods for mixture models. Since this problem has long been a central interest of mine’s and since the method is closely connected with the harmonic mean solution we developed with Darren Wraith in 2009, (and also included in our 2009 survey with Jean-Michel Marin of evidence approximations, published in Frontiers of Statistical Decision Making and Bayesian Analysis for Jim Berger’s 60th birthday), I quickly went into the paper. The core purpose of this paper is to adapt THAMES to a multimodal setting since using an ellipsoidal region as the support of the Uniform reciprocal importance sampling distribution does not make sense for a multimodal target. After reading it a few times, and while some computational aspects remain obscure to me, I am not convinced this brings an adequate answer to the challenge.  Indeed, while the approach borrows directly from Berkhof et al. (2003) that inspired the resolution we proposed, Jeong (Kate) Lee and myself, the issues I have with the current proposal are that

1. the evacuation of earlier methods as not simple or not universal enough is rather disingenuous. For instance, software that do not return (latent) allocation vectors can easily be post-processed. And the current method uses allocation probabilities just the same (in Section 3.3). Similarly, the random shuffling answer to label (lack of) switching proposed by Sylvia Früwirth-Schnatter—which again can be achieved by post-processing—cannot be rejected on the sole basis that the component means (based on the MCMC sample) are all similar. It is furthermore debatable that the current proposal is simple, when involving relabelling à la Stephens, averaging over permutations, selecting over said permutations by constructing a graph over components (section 3.2.1) and  running a quadratic discriminant analysis (section 3.2.2) on the posterior sample, based on an arbitrary Normal representation of the distributions of the clusters, and finally defining a new ordering constraint (section 3.2.3). Computing efforts  required by the respective methods do not appear in the main text.

2. the handling of the label switching issue—the reason why Larry Wasserman saw mixtures at the same magnitude of evil as tequila!—is problematic for several reasons. My position (since at least 2000!) on the matter is that the proper posterior sample must exhibit label switching and come close to symmetry among the “components”. The label switching problem (section 3.1) is rather when the MCMC sample does not “switch their labels”. The relabelling approach (e.g., à la Stephens) allows for a differentiation between components, to some extent, which helps with computing basic posterior moments for point estimation or for the calibration of the support of the Uniform reciprocal importance sampling distribution, but the use of any relabelling procedure is tampering with the original MCMC sample and thus bound to impact the distribution of the resulting relabelled sample. Furthermore, relabelling depends on the value of G, whereas the actual number of (significant) modes in the posterior is also connected with the (partial) fit of the data to the model, meaning the creation of further modes than those linked with relabelling. Especially when the model is misspecified. Incidentally, the symmetrised version of THAMES (5) does not require relabelling. Neither does the Bayes factor. In addition, the experiment section (4.1.2) mentions that bridge sampling is biased by a factor of G!, which comes as a surprise to me since I associated this factor with the call to Sid Chib’s formula in the absence of label switching, i.e. when the MCMC sample was stuck on a mode, as exposed by Radford Neal in 1999. Is it because bridge sampling is applied to the relabelled sample? It is also surprising that the gap appears in the simulated datasets (Fig.3) and not in the real ones (Fig.5).

3. the (legitimate) purpose of using marginal likelihoods for selecting the number G of components is weakened by the intrusion of alternate proposals to assess G from the data, like the criterion of overlap (section 3.2.1), which instead aims at the number of clusters, with an elimination of “empty components” that should either remain a possibility (within a regular mixture model) or be evacuated with a different modelling (à la Diebolt & Robert, or à la Wasserman). This overlapping criterion is further used in the discriminant analysis that only applies to “non-overlapping components” of the mixture (section 3.2.3)—at which point I got lost in the reordering and simplification of the computation of THAMES (but got reminded of the results of Agostino Nobile in the 2000’s, with whom I used to discuss a lot in my yearly visit to the University of Glasgow).

4. several mentions are made of the other estimators being biased, which is indeed the case for bridge sampling (if not necessarily for importance sampling), but not necessarily a central issue, while the original generalised harmonic proposal by Gelfand and Dey (1994) and thus THAMES produce an unbiased estimator of the inverse of the evidence (thus neither of the evidence nor of the log-evidence). However, in the paper, the volume of the support of the Uniform reciprocal importance sampling distribution is estimated by a basic Monte Carlo coverage probability in (3), which induces the same type of bias as the other methods.

miXtures on arXiv

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , on February 5, 2025 by xi'an

A paper about Bayesian inference on mixtures was posted on arXiv last week, as of 13 Jan 2025.  Fast sampling and model selection for Bayesian mixture models, by M. E. J. Newman is based on the notion that (genuine) parameters of a mixture model can be marginalized out when using conjugate priors. This is something that we pointed out quite a while ago, in a 1999 paper with George and Marty, which was devised in a long ride from Baltimore to Cornell after JSM 1999, and again in the 2002 Series B perfect sampling paper with George, Kerrie and Mike. (Also written in 1999.) And marginal likelihood can furthermore be approximated along this way as discussed in the more recent papers Bayesian Inference on Mixtures of Distributions with Kate, Kerrie & Jean-Michel, as well as Approximating the marginal likelihood in mixture models with Jean-Michel.

“Standard mixture models, as commonly formulated, also suffer from a technical, but important, difficulty: the existence of empty components. In many models (…) the number of observations in a component can be zero. Arguably this is acceptable for a model with a fixed number of components, but when the number of components is a free random variable it causes ambiguity, because a given division of observations into components can be represented in more than one way in the model. For instance, we could divide observations into two components, or we could divide them into three components, one of which is empty. This in turn creates difficulties when estimating the number of components—do we have two components or three?”

A very puzzling perspective, imho, since potentially empty components are inherent to (both finite and infinite) mixture models with connected issues of prohibiting some improper priors (if not all) and non-identifiability, including non-identifiability of the number of empty components (which remains random conditional on the data!), but different numbers of components lead to different models and their comparison is handled straightforwardly by a Bayesian analysis.

The author then proceeds to “prohibit empty components” [as a prior choice ?] as we did in the original (!) Gibbs sampler for mixtures in 1990 (published in 1994 in Series B!), seeking posterior properness, a trick later validated by Larry Wasserman (in again 1999, the year of mixtures!). Who called the construct the combination of a fixed prior and of a pseudo-likelihood, correctly imho (as the data dependent part is not properly normalised by a function of the parameters), rather than a prior choice. (The very one who stated that “mixtures, like tequila, are evil and should be avoided“.)

From there, the modelling is rather standard, with an arbitrary prior on k, number of components, a random partition model that prohibits empty components, even though the constraint could be more stringent depending on the number of parameters of a given component and the degree of improperness of the prior, as in our 1990 Series B paper. (Impropriety is not discussed in the paper.) Bayesian inference on k is based on the simulated (pseudo-)posterior. The choice therein as the estimated clustering is the most frequent partition (consensus clustering), connected to our proposal of (again!) 1999 with Merrilee and Gilles. While the estimated mixture is not explicited. The approach is assessed as running at an O(k) cost, with no parallel in terms of the data size n, even though the examples include a 59,946 dataset. One notable algorithmic trick when moving k is in selecting a component at random first rather than an observation index.

Some minor issues: detailed balance indicated as required for convergence (p14), label switching is called component switching (p5), higher acceptance rate indicated as meaning improved performances (p7)

Bill’s 80th!!!

Posted in pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , on April 17, 2022 by xi'an

“It was the best of times,
it was the worst of times”
[Dickens’ Tale of Two Cities (which plays a role in my friendship with Bill!)]

My flight to NYC last week was uneventful and rather fast and I worked rather well, even though the seat in front of me was inclined to the max for the entire flight! (Still got glimpses of Aline and of Deepwater Horizon from my neighbours.) Taking a very early flight from Paris was great making a full day once in NYC,  but “forcing” me to take a taxi, which almost ended up in disaster since the Über driver did not show up. At all. And never replied to my message. Fortunately trains were running, I was also running despite the broken rib, and I arrived at the airport some time before access was closed, grateful for the low activity that day. I also had another bit of a worrying moment at the US border control in JFK as I ended up in a back-office of the Border Police after the machine could not catch my fingerprints. And another stop at the luggage control as my lack of luggage sounded suspicious!The conference was delightful in celebrating Bill’s carreer and kindness (tinted with the most gentle irony!). Among stories told at the banquet, I was surprised to learn of Bill’s jazz career side, as I had never heard him play the piano or the clarinet! Even though we had chatted about music and literature on many occasions. Since our meeting in 1989… The (scientific side of the) conference included many talks around shrinkage, from loss estimation to predictive estimation, reminding me of the roaring 70’s and 80’s [James-Stein wise]. And demonstrating the impact of Bill’s wor throughout this era (incl. on my own PhD thesis). I started wondering at the (Bayesian) use of the loss estimate, though, as I set myself facing two point estimators attached with two estimators of their loss: it did not seem a particularly good idea to systematically pick the one with the smallest estimate (and Jim Berger confirmed this feeling on a later discussion). Among the talks on less familiar topics (of mine), I discovered work of Genevera Allen‘s on inferring massive network for neuron connections under sparse information. And of Emma Jingfei Zhang, equally centred on network inference, with applications to brain connectivity.

In a somewhat remote connection with Bill’s work (and our joint and hilarious assessment of Pitman closeness), I presented part of our joint and current work with Adrien Hairault and Judith Rousseau on inferring the number of components in a mixture by Bayes factors when the alternative is an infinite mixture (i.e., a Dirichlet process mixture). Of which Ruobin Gong gave a terrific discussion. (With a connection to her current work on Sense and Sensitivity.)

I was most sorry to miss Larry Wasserman’s and Rob Strawderman’s talk to rush back to the airport, the more because I am sure Larry’s talk would have brought a new light on causality (possibly equating it with tequila and mixtures!). The flight back was uneventfull, the plane rather empty and I slept most of the time. Overall,  it was most wonderful to re-connect with so many friends. Most of whom I had not seen for ages, even before the pandemic. And to meet new friends. (Nothing original in the reported feeling, just telling that the break in conferences and workshops was primarily a hatchet job on social relations and friendships.)

repulsive mixtures

Posted in Books, Statistics with tags , , , , , , , , on April 10, 2017 by xi'an

Fangzheng Xie and Yanxun Xu arXived today a paper on Bayesian repulsive modelling for mixtures. Not that Bayesian modelling is repulsive in any psychological sense, but rather that the components of the mixture are repulsive one against another. The device towards this repulsiveness is to add a penalty term to the original prior such that close means are penalised. (In the spirit of the sugar loaf with water drops represented on the cover of Bayesian Choice that we used in our pinball sampler, repulsiveness being there on the particles of a simulated sample and not on components.) Which means a prior assumption that close covariance matrices are of lesser importance. An interrogation I have has is was why empty components are not excluded as well, but this does not make too much sense in the Dirichlet process formulation of the current paper. And in the finite mixture version the Dirichlet prior on the weights has coefficients less than one.

The paper establishes consistency results for such repulsive priors, both for estimating the distribution itself and the number of components, K, under a collection of assumptions on the distribution, prior, and repulsiveness factors. While I have no mathematical issue with such results, I always wonder at their relevance for a given finite sample from a finite mixture in that they give an impression that the number of components is a perfectly estimable quantity, which it is not (in my opinion!) because of the fluid nature of mixture components and therefore the inevitable impact of prior modelling. (As Larry Wasserman would pound in, mixtures like tequila are evil and should likewise be avoided!)

The implementation of this modelling goes through a “block-collapsed” Gibbs sampler that exploits the latent variable representation (as in our early mixture paper with Jean Diebolt). Which includes the Old Faithful data as an illustration (for which a submission of ours was recently rejected for using too old datasets). And use the logarithm of the conditional predictive ordinate as  an assessment tool, which is a posterior predictive estimated by MCMC, using the data a second time for the fit.