Archive for Charlie Geyer

multimodal challenges

Posted in Books, Statistics, University life with tags , , , , , , , , , , on April 30, 2026 by xi'an

At the last mostly Monte Carlo seminar, Pierre Monmarché presented a recent work on post-sampling for multimodal targets: while I  consider the main problem in sampling from generic multimodal targets stands with finding the modes, rather than with exploring local aspects or estimating relative weights of said modes, this made me ponder whether or not this could be accelerated by removing chunks of the already explored modes to induce moves elsewhere, which is a form of radical, brute-force, tempering, or of Wang-Landau.  As for the relative weights, a multiple move proposal can be considered, including our folding idea. Or Geyer’s inverse logistic trick. Or the similar mixture trick we used in our Biometrika paper on nested sampling. Pierre’s approach was closer to adaptive importance sampling, with a self-imposed constraint of fixed sample sizes from (approximate) distributions around each of the modes.

robust simulation-based inference

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , on March 7, 2026 by xi'an

This new arXival by Lorenzo Tomaselli, Valérie Ventura, and Larry Wasserman (from CMU) considers simulation-based inference under model misspecification (as we did for ABC in our 2020 Series B paper). Which is almost always the case. In the paper, SBI is defined as producing N parameters and N samples from the prior and the corresponding sampling distribution, respectively, and then doubling the resulting samples by permuting at random the parameters θ. This means that the second half is distributed from the product of the prior and of the marginal, hence that the classification odds ratio is equal to the likelihood, hence providing an estimation method (andlikelihood trick) à la Geyer. From this estimate, an ABC p-value can be derived, but it is incorrect as such when the model is misspecified. Hence the use of the Hellinger discrepancy, the power divergence and the kernel distance (or MMD) as alternatives to the misspecified MLE.

The paper then expands on approximating density ratios by virtue of a reproducing kernel Hilbert space, using a Gaussian kernel. (With a nice remark on requiring only one single ratio estimator for all values of θ, albeit in the joint space.) And focus on a studentized MMD estimator (à la e-value) to build a confidence set that remains valid under model misspecification. And without regularity assumptions.

Another approach is further explored, based on exponential tilting—of which I am not a great fan, from being highly dependent on the choice of the pseudo-sufficient statistic to require an intractable normalising constant, to requiring an extra optimization, even though I appreciate the mathematical appeal of the construct. Which seems to require a sample simulation for each value of θ at the learning stage, albeit relying on the same likelihood trick. The appropriateness of the tilting can be tested by a goodness of fit test tailored for the SBI structure, which sounds rather greedy in the required simulations. 

Besides the g-and-k distribution example (which, as pointed out several times on the ‘Og, is not intractable, strictly speaking!), the paper studies a mixture example, despite Larry dubbing them as evil as tequila a long while ago! (The paper also offers a section called accoutrements, which is my first encounter with this use of the term, usually found in medieval contexts!)

Note that Larry will present the paper at the OWABI webinar next 25 March!

deep Bayes factor

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , on August 8, 2024 by xi'an

A recently arXived paper proposes an alternative approach to computing Bayes factors via deep learning, Deep Bayes Factors written by Jungeum Kim (presenting her work at JSM this very morning) and Veronika Ročková (whom I have known from her PhD years and whose COPSS Award we very gladly celebrated yesterday!). Which is obviously of interest to me, given my repeated visits to the challenge.

“we introduce Deep Bayes Factor (DeepBF), a neural classifier trained on simulated datasets to learn a mapping whose functional constitutes a Bayes factor estimator.”

Their approach is directly connected with various classification approaches to ABF, incl. the mythical inverse logistic version of Geyer (1994) and noise contrastive estimation of Gutmann and Hyvärinen (2010) (as well as our forested version). Which is called the likelihood-ratio trick here.

“Viewing the Bayes factor through the lens of binary classification aligns with Pudlo et al. (2016), who recast ABC model selection as a classification problem. They employ random forests to select a model by a majority vote. Instead, we focus on binary classification where the purpose is to learn marginal likelihood ratios.  Contrary to the method in Pudlo et al. (2016), our strategy circumvents a secondary learning phase for gauging model posterior estimates, delivering results in only one stage.”

The authors‘ solution stands with learning a classifier from simulated data from both models (and a basic log ratio utility), along iterations updating D from the gradient of the utility, the associated Bayes factor being the ratio D/(1-D) derived from the estimated classifier. There is a cost in producing new samples from the (same) predictives at each iteration (and I wonder if some recycling would be helpful, as well as reducing the sample size for the simpler model). In one of the remarks, the authors point out that “in the effort to see the best ABC performance, we intentionally use the full data Y as a summary statistic”, a remark that I find surprising given the overall consensus that the Bayes factor itself [when based on the full data] is close to optimal.

The method is overall consistent (in the data size n) under classical Bayesian asymptotics, sometimes even when the Bayes factor estimator is inconsistent, naturally expands to pseudo Bayes factors like intrinsic and fractional Bayes factors, also mileage varies in terms of numerical stability.

In the Bayesian model criticism section, the notion of opposing the actual dataset to a simulated one relates very much to Geyer’s (1994) solution. As well as to GANs, as noted in the paper. I did not look closely at the numerical comparisons in the experimental section, but they sound rich enough.

evidence estimation in finite and infinite mixture models

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , on May 20, 2022 by xi'an

Adrien Hairault (PhD student at Dauphine), Judith and I just arXived a new paper on evidence estimation for mixtures. This may sound like a well-trodden path that I have repeatedly explored in the past, but methinks that estimating the model evidence doth remain a notoriously difficult task for large sample or many component finite mixtures and even more for “infinite” mixture models corresponding to a Dirichlet process. When considering different Monte Carlo techniques advocated in the past, like Chib’s (1995) method, SMC, or bridge sampling, they exhibit a range of performances, in terms of computing time… One novel (?) approach in the paper is to write Chib’s (1995) identity for partitions rather than parameters as (a) it bypasses the label switching issue (as we already noted in Hurn et al., 2000), another one is to exploit  Geyer (1991-1994) reverse logistic regression technique in the more challenging Dirichlet mixture setting, and yet another one a sequential importance sampling solution à la  Kong et al. (1994), as also noticed by Carvalho et al. (2010). [We did not cover nested sampling as it quickly becomes onerous.]

Applications are numerous. In particular, testing for the number of components in a finite mixture model or against the fit of a finite mixture model for a given dataset has long been and still is an issue of much interest and diverging opinions, albeit yet missing a fully satisfactory resolution. Using a Bayes factor to find the right number of components K in a finite mixture model is known to provide a consistent procedure. We furthermore establish there the consistence of the Bayes factor when comparing a parametric family of finite mixtures against the nonparametric ‘strongly identifiable’ Dirichlet Process Mixture (DPM) model.

accronyms [CDT lectures]

Posted in Books, Statistics with tags , , , , , , , , , , , , , , , on May 16, 2022 by xi'an

This week, I gave a short and introductory course in Warwick for the CDT (PhD) students on my perceived connections between reverse logistic regression à la Geyer and GANS, among other things. The first attempt was cancelled in 2020 due to the pandemic, the second one in 2021 was on-line and thus offered little possibilities for interactions. Preparing for this third attempt made me read more papers on some statistical analyses of GANs and WGANs, which was more satisfactory [for me] even though I could not get into the technical details…