Archive for marginalisation

miXtures on arXiv

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , on February 5, 2025 by xi'an

A paper about Bayesian inference on mixtures was posted on arXiv last week, as of 13 Jan 2025.  Fast sampling and model selection for Bayesian mixture models, by M. E. J. Newman is based on the notion that (genuine) parameters of a mixture model can be marginalized out when using conjugate priors. This is something that we pointed out quite a while ago, in a 1999 paper with George and Marty, which was devised in a long ride from Baltimore to Cornell after JSM 1999, and again in the 2002 Series B perfect sampling paper with George, Kerrie and Mike. (Also written in 1999.) And marginal likelihood can furthermore be approximated along this way as discussed in the more recent papers Bayesian Inference on Mixtures of Distributions with Kate, Kerrie & Jean-Michel, as well as Approximating the marginal likelihood in mixture models with Jean-Michel.

“Standard mixture models, as commonly formulated, also suffer from a technical, but important, difficulty: the existence of empty components. In many models (…) the number of observations in a component can be zero. Arguably this is acceptable for a model with a fixed number of components, but when the number of components is a free random variable it causes ambiguity, because a given division of observations into components can be represented in more than one way in the model. For instance, we could divide observations into two components, or we could divide them into three components, one of which is empty. This in turn creates difficulties when estimating the number of components—do we have two components or three?”

A very puzzling perspective, imho, since potentially empty components are inherent to (both finite and infinite) mixture models with connected issues of prohibiting some improper priors (if not all) and non-identifiability, including non-identifiability of the number of empty components (which remains random conditional on the data!), but different numbers of components lead to different models and their comparison is handled straightforwardly by a Bayesian analysis.

The author then proceeds to “prohibit empty components” [as a prior choice ?] as we did in the original (!) Gibbs sampler for mixtures in 1990 (published in 1994 in Series B!), seeking posterior properness, a trick later validated by Larry Wasserman (in again 1999, the year of mixtures!). Who called the construct the combination of a fixed prior and of a pseudo-likelihood, correctly imho (as the data dependent part is not properly normalised by a function of the parameters), rather than a prior choice. (The very one who stated that “mixtures, like tequila, are evil and should be avoided“.)

From there, the modelling is rather standard, with an arbitrary prior on k, number of components, a random partition model that prohibits empty components, even though the constraint could be more stringent depending on the number of parameters of a given component and the degree of improperness of the prior, as in our 1990 Series B paper. (Impropriety is not discussed in the paper.) Bayesian inference on k is based on the simulated (pseudo-)posterior. The choice therein as the estimated clustering is the most frequent partition (consensus clustering), connected to our proposal of (again!) 1999 with Merrilee and Gilles. While the estimated mixture is not explicited. The approach is assessed as running at an O(k) cost, with no parallel in terms of the data size n, even though the examples include a 59,946 dataset. One notable algorithmic trick when moving k is in selecting a component at random first rather than an observation index.

Some minor issues: detailed balance indicated as required for convergence (p14), label switching is called component switching (p5), higher acceptance rate indicated as meaning improved performances (p7)

Saddlepoint Monte Carlo and its application to vote transfers

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , on January 3, 2024 by xi'an


Our former Dauphine Master student Théo Voldoire (now PhD-ing at Harvard), along with Nicolas Chopin, Guillaume Rateau, and Robin Ryder (now at Imperial), arXived a month ago a paper with title Saddlepoint Monte Carlo and its Application to Exact Ecological Inference, essentially the outcome of his Master thesis last summer. Nicolas came to present the paper at our Mostly MC seminar. The motivating example is about vote transfers (to surviving candidates for a second round, as in the French presidential and deputorial elections) based on results from polling stations across rounds, which equates filling a contingency table with known margins, exploiting multiple tools like characteristic functions, inverse Fourier transform, its Monte Carlo version, pseudo-marginal MCMC, tilting, exponential families, quasi Monte Carlo! Which also reminded me of the time Reuven Rubinstein was occasionally visiting Paris, defending the cross entropy approach. Among multiple questions raised by this original approach to an “old” problem, one may think of the model misspecification issue that political analysts would not fail to raise, namely that the transfer estimates are based on multinomial models, that all models are wrong, &tc. We discussed briefly about this during the seminar, the suggestion being a predictive check by cross validation. The talk also brought to mind highly probable applications to privacy, and possibly to capture recapture.

individualised polychotomous logistic regression

Posted in Books, Statistics, University life with tags , , , , , on May 17, 2022 by xi'an

A recent submission to Biometrika made me read the 1984 Biometrika paper of Begg and Gray on the individualisation of polychotomous regression, namely the idea that when considering this model with T categories, the regression parameters could be estimated by considering only the pairs (0,i), 0 being the baseline category (with no parameter), since the (true) probability to be in category i conditional on being in either category 0 or category i is logistic with the same coefficient as in the polychotomous model. While I see no issue with this remark (contrary to the submission author), it is of course producing (much quicker) a different estimate of the polychotomous parameter, when compared with the full likelihood approach. Not only because it does not exploit the entire information contained in the data but also because it operates with a pseudo-likelihood.

efficiency of normalising over discrete parameters

Posted in Statistics with tags , , , , , , , , , on May 1, 2022 by xi'an

Yesterday, I noticed a new arXival entitled Investigating the efficiency of marginalising over discrete parameters in Bayesian computations written by Wen Wang and coauthors. The paper is actually comparing the simulation of a Gibbs sampler with an Hamiltonian Monte Carlo approach on Gaussian mixtures, when including and excluding latent variables, respectively. The authors missed the opposite marginalisation when the parameters are integrated.

While marginalisation requires substantial mathematical effort, folk wisdom in the Stan community suggests that fitting models with marginalisation is more efficient than using Gibbs sampling.

The comparison is purely experimental, though, which means it depends on the simulated data, the sample size, the prior selection, and of course the chosen algorithms. It also involves the [mostly] automated [off-the-shelf] choices made in the adopted software, JAGS and Stan. The outcome is only evaluated through ESS and the (old) R statistic. Which all depend on the parameterisation. But evacuates the label switching problem by imposing an ordering on the Gaussian means, which may have a different impact on marginalised and unmarginalised models. All in all, there is not much one can conclude about this experiment since the parameter values beyond the simulated data seem to impact the performances much more than the type of algorithm one implements.

a case for Bayesian deep learnin

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , on September 30, 2020 by xi'an

Andrew Wilson wrote a piece about Bayesian deep learning last winter. Which I just read. It starts with the (posterior) predictive distribution being the core of Bayesian model evaluation or of model (epistemic) uncertainty.

“On the other hand, a flat prior may have a major effect on marginalization.”

Interesting sentence, as, from my viewpoint, using a flat prior is a no-no when running model evaluation since the marginal likelihood (or evidence) is no longer a probability density. (Check Lindley-Jeffreys’ paradox in this tribune.) The author then goes for an argument in favour of a Bayesian approach to deep neural networks for the reason that data cannot be informative on every parameter in the network, which should then be integrated out wrt a prior. He also draws a parallel between deep ensemble learning, where random initialisations produce different fits, with posterior distributions, although the equivalent to the prior distribution in an optimisation exercise is somewhat vague.

“…we do not need samples from a posterior, or even a faithful approximation to the posterior. We need to evaluate the posterior in places that will make the greatest contributions to the [posterior predictive].”

The paper also contains an interesting point distinguishing between priors over parameters and priors over functions, ony the later mattering for prediction. Which must be structured enough to compensate for the lack of data information about most aspects of the functions. The paper further discusses uninformative priors (over the parameters) in the O’Bayes sense as a default way to select priors. It is however unclear to me how this discussion accounts for the problems met in high dimensions by standard uninformative solutions. More aggressively penalising priors may be needed, as those found in high dimension variable selection. As in e.g. the 10⁷ dimensional space mentioned in the paper. Interesting read all in all!