Archive for marginal likelihood identity

THAMES for mixtures, a reply from the authors

Posted in Books, pictures, R, Statistics, University life with tags , , , , , , , , , , , , , , on June 23, 2025 by xi'an

[Here is a reply to my comments on THAMES sent by the first author of the paper, Martin Metodiev. The above replica of the cover of Rivers of London is obviously unrelated with the reply or the original blog, beyond presenting a fantasy map of the Thames!]

Thank you for your review of our article! Adapting your previous work in this field has been a pleasure. Before I respond to your comments, I would like to emphasize that the simplicity of our estimator lies in its simple analytic expression (a truncated harmonic mean of reciprocal unnormalized posterior density values). Indeed, our package “thamesmix” (recently submitted to CRAN!) has a function to compute the marginal likelihood of any mixture model. This function requires only two parameters: the unnormalized log-posterior function (the logarithm of the prior plus the log-likelihood) and the MCMC simulations from the posterior.

Regarding your main comments:

1. “the evacuation of earlier methods as not simple or not universal enough is rather disingenuous. For instance, software that do not return (latent) allocation vectors can easily be post-processed.”

I could not find an example of post-process simulations on top of MCMC outputs applied to compute these methods. It sounds really interesting, and I would be happy to cite it. Is there a reference that you can recommend?

In any case, the point still stands. Most estimators which we cite with regards to this point do not just need allocation samplers, but also the analytic expressions of the distribution of the allocation vectors or the distribution of the data conditional on these allocation vectors that come with them. I do not think that a closed form of this distribution is available in general.

2.“the handling of the label switching issue—the reason why Larry Wasserman saw mixtures at the same magnitude of evil as tequila!—is problematic for several reasons.”

The fact that our estimator is invariant to label-switching is indeed the core of our method. The simple Gibbs sampler gets stuck in one mode, and this is why the classical version of bridge sampling is biased by a factor of G! in the simulation setting. As you point out, this is successfully resolved when using fully symmetric bridge sampling in the experiment section. However, the computation cost of this fully symmetric estimator rises super-exponentially with G, so I do not see how it could be evaluated for G=15, where the number of symmetric modes is equal to 15! (over one trillion). One of the main points of our article is that the symmetric THAMES can be evaluated in a feasible amount of time, even in such a high-dimensional multivariate setting.

3. “the (legitimate) purpose of using marginal likelihoods for selecting the number G of components is weakened by the intrusion of alternate proposals to assess G from the data”

I would like to point out that these alternate proposals do not in any way impact the definition of the THAMES. It is the simple definition given in Equation (5). They are only used to speed up the computation.

4. “several mentions are made of the other estimators being biased, which is indeed the case for bridge sampling (if not necessarily for importance sampling), but not necessarily a central issue”

The problem that we see with the classical, non-symmetric bridge sampling method in the setting of mixture models is not simply that it is biased. The problem is that the bias is persistent and often roughly equal to the factor of G! when the MCMC sampler failed to switch between modes. We have not had this experience with the THAMES: it converged even when the MCMC was stuck.

variational approximation to empirical likelihood ABC

Posted in Statistics with tags , , , , , , , , , , , , , , , , , , on October 1, 2021 by xi'an

Sanjay Chaudhuri and his colleagues from Singapore arXived last year a paper on a novel version of empirical likelihood ABC that I hadn’t yet found time to read. This proposal connects with our own, published with Kerrie Mengersen and Pierre Pudlo in 2013 in PNAS. It is presented as an attempt at approximating the posterior distribution based on a vector of (summary) statistics, the variational approximation (or information projection) appearing in the construction of the sampling distribution of the observed summary. (Along with a weird eyed-g symbol! I checked inside the original LaTeX file and it happens to be a mathbbmtt g, that is, the typewriter version of a blackboard computer modern g…) Which writes as an entropic correction of the true posterior distribution (in Theorem 1).

“First, the true log-joint density of the observed summary, the summaries of the i.i.d. replicates and the parameter have to be estimated. Second, we need to estimate the expectation of the above log-joint density with respect to the distribution of the data generating process. Finally, the differential entropy of the data generating density needs to be estimated from the m replicates…”

The density of the observed summary is estimated by empirical likelihood, but I do not understand the reasoning behind the moment condition used in this empirical likelihood. Indeed the moment made of the difference between the observed summaries and the observed ones is zero iff the true value of the parameter is used in the simulation. I also fail to understand the connection with our SAME procedure (Doucet, Godsill & X, 2002), in that the empirical likelihood is based on a sample made of pairs (observed,generated) where the observed part is repeated m times, indeed, but not with the intent of approximating a marginal likelihood estimator… The notion of using the actual data instead of the true expectation (i.e. as a unbiased estimator) at the true parameter value is appealing as it avoids specifying the exact (or analytical) value of this expectation (as in our approach), but I am missing the justification for the extension to any parameter value. Unless one uses an ancillary statistic, which does not sound pertinent… The differential entropy is estimated by a Kozachenko-Leonenko estimator implying k-nearest neighbours.

“The proposed empirical likelihood estimates weights by matching the moments of g(X¹), …, g(X⁹) with that of
g(X⁰), without requiring a direct relationship with the parameter. (…) the constraints used in the construction of the empirical likelihood are based on the identity in (7), which can only be satisfied when θ = θ⁰. “

Although I am feeling like missing one argument, the later part of the paper seems to comfort my impression, as quoted above. Meaning that the approximation will fare well only in the vicinity of the true parameter. Which makes it untrustworthy for model choice purposes, I believe. (The paper uses the g-and-k benchmark without exploiting Pierre Jacob’s package that allows for exact MCMC implementation.)

back to Ockham’s razor

Posted in Statistics with tags , , , , , , , , , on July 31, 2019 by xi'an

“All in all, the Bayesian argument for selecting the MAP model as the single ‘best’ model is suggestive but not compelling.”

Last month, Jonty Rougier and Carey Priebe arXived a paper on Ockham’s factor, with a generalisation of a prior distribution acting as a regulariser, R(θ). Calling on the late David MacKay to argue that the evidence involves the correct penalising factor although they acknowledge that his central argument is not absolutely convincing, being based on a first-order Laplace approximation to the posterior distribution and hence “dubious”. The current approach stems from the candidate’s formula that is already at the core of Sid Chib’s method. The log evidence then decomposes as the sum of the maximum log-likelihood minus the log of the posterior-to-prior ratio at the MAP estimator. Called the flexibility.

“Defining model complexity as flexibility unifies the Bayesian and Frequentist justifications for selecting a single model by maximizing the evidence.”

While they bring forward rational arguments to consider this as a measure model complexity, it remains at an informal level in that other functions of this ratio could be used as well. This is especially hard to accept by non-Bayesians in that it (seriously) depends on the choice of the prior distribution, as all transforms of the evidence would. I am thus skeptical about the reception of the argument by frequentists…