Archive for the R Category
against modern sl[AI]very
Posted in Kids, R, Statistics, University life with tags AI, AIMS, artificial intelligence, data mining, ethical machine learning, hackathon, Human Rights, IRC, LLM, machine learning research, modern slavery, open source, Paso Libres, QUT, software development, UNESCO on July 2, 2025 by xi'anTHAMES for mixtures, a reply from the authors
Posted in Books, pictures, R, Statistics, University life with tags allocations, Ben Aaronovitch, bridge sampling, CRAN, harmonic mean estimator, label switching, London, marginal likelihood, marginal likelihood identity, MCMC, R, response, Rivers of London, Thames, unbiasedness on June 23, 2025 by xi'an
[Here is a reply to my comments on THAMES sent by the first author of the paper, Martin Metodiev. The above replica of the cover of Rivers of London is obviously unrelated with the reply or the original blog, beyond presenting a fantasy map of the Thames!]
Thank you for your review of our article! Adapting your previous work in this field has been a pleasure. Before I respond to your comments, I would like to emphasize that the simplicity of our estimator lies in its simple analytic expression (a truncated harmonic mean of reciprocal unnormalized posterior density values). Indeed, our package “thamesmix” (recently submitted to CRAN!) has a function to compute the marginal likelihood of any mixture model. This function requires only two parameters: the unnormalized log-posterior function (the logarithm of the prior plus the log-likelihood) and the MCMC simulations from the posterior.
Regarding your main comments:
1. “the evacuation of earlier methods as not simple or not universal enough is rather disingenuous. For instance, software that do not return (latent) allocation vectors can easily be post-processed.”
I could not find an example of post-process simulations on top of MCMC outputs applied to compute these methods. It sounds really interesting, and I would be happy to cite it. Is there a reference that you can recommend?
In any case, the point still stands. Most estimators which we cite with regards to this point do not just need allocation samplers, but also the analytic expressions of the distribution of the allocation vectors or the distribution of the data conditional on these allocation vectors that come with them. I do not think that a closed form of this distribution is available in general.
2.“the handling of the label switching issue—the reason why Larry Wasserman saw mixtures at the same magnitude of evil as tequila!—is problematic for several reasons.”
The fact that our estimator is invariant to label-switching is indeed the core of our method. The simple Gibbs sampler gets stuck in one mode, and this is why the classical version of bridge sampling is biased by a factor of G! in the simulation setting. As you point out, this is successfully resolved when using fully symmetric bridge sampling in the experiment section. However, the computation cost of this fully symmetric estimator rises super-exponentially with G, so I do not see how it could be evaluated for G=15, where the number of symmetric modes is equal to 15! (over one trillion). One of the main points of our article is that the symmetric THAMES can be evaluated in a feasible amount of time, even in such a high-dimensional multivariate setting.
3. “the (legitimate) purpose of using marginal likelihoods for selecting the number G of components is weakened by the intrusion of alternate proposals to assess G from the data”
I would like to point out that these alternate proposals do not in any way impact the definition of the THAMES. It is the simple definition given in Equation (5). They are only used to speed up the computation.
4. “several mentions are made of the other estimators being biased, which is indeed the case for bridge sampling (if not necessarily for importance sampling), but not necessarily a central issue”
The problem that we see with the classical, non-symmetric bridge sampling method in the setting of mixture models is not simply that it is biased. The problem is that the bias is persistent and often roughly equal to the factor of G! when the MCMC sampler failed to switch between modes. We have not had this experience with the THAMES: it converged even when the MCMC was stuck.
proximal sampler
Posted in Books, pictures, R, Statistics, Travel, University life with tags banana, Bayesian optimization, Columbia University, log-concave functions, Metropolis-Hastings algorithm, New York, New York city, proximal sampler, random walk Metropolis algorithm, Statistical learning, ULA, unadjusted Langevin algorithm, workshop on April 28, 2025 by xi'an
At the Columbia workshop last week, Andre Wibisono presented work related with a recent arXival on the exponentially fast convergence of both unadjusted Langevin and proximal sampler algorithms under strong [definitely strong] log-concavity assumptions. The idea behind the proximal sampler is to target the demarginalised density
by introducing an auxiliary Gaussian vector y, which preserves f(x) as the marginal distribution on the first component vector X. While the auxiliary Y is (obviously) conditionally Gaussian, the conditional of X is at least as challenging as simulating from f. Unless η is chosen small enough to regularize log g(x) into a strongly log-concave function, since
when x*=x*(y) is the maximiser of log g(x,y) (for a given value y) and β>0 is the appropriate log-concavity constant. This inequality means that an accept-reject can be implemented to simulate from the conditional of X given Y but it requires both the factor β and the derivation of x*(y), hence a pretty good understanding and a rather high regularity of the actual target f(x). Besides, the regularization term ||x-y||² means that y is approximately the previous value of the (sub)chain X, hence it creates a rappel force that slows down the exploration of the target.
Since the arXival does not contain numerical comparisons, I attempted one using the (2D) banana shaped distribution,
target=function(x,sig,B,mu)-x[1]^2/2/sig-(x[2]+B*x[1]^2-mu)^2/2
with μ=σ=B=10. Comparing with a vanilla random walk Metropolis with three potential scales, chosen randomly at each iteration. Since I did not want to check whether or not the target was log-concave (and derive the corresponding β), I used the Normal distribution centred at proposal x*(y) of a Metropolis step, again with several scales. The following is the representation of the samples (sienna for MCMC, navy blue for proximal with β=50, dark green for β=5), with a lesser rate of tail exploration for the proximal samplers. It is thus unclear to me the theoretical characterisations of the method translate into practical efficiency beyond the most regular cases.
computo on the go
Posted in Books, R, Statistics, University life with tags binary, Biometrika, Computo, editor in chief, free software, French researchers, github, journal, Julia, Latin, logo, machine learning, open access, open source, Python, Quarto, R, repositories, reproducible research, Rmarkdown, SFDS, Société française de Statistique, Statistics on April 8, 2025 by xi'an

