Archive for identifiability

mixture models [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , on August 14, 2024 by xi'an

Strangely enough, I became aware of this new book on mixtures through one of these annoying emails “Your work has been cited n times this week“… Mixture Models (Parametric, Semiparametric, and New Directions) by Weixin Yao and Sijia Wang got published by CRC Press earlier this year, within the Monographs on Statistics and Applied Probability green series (#175), and covers across 380 pages most aspects of mixture (and hidden Markov) estimation, if with strong emphasis on maximum likelihood estimation, while the new directions are unsurprisingly those pursued by the authors, namely robust and semi-parametric estimation, as well as model selection by testing.

An early warning about this book review is that I co-edited a Handbook of Mixture Analysis with my friends Sylvia Früwirth-Schnatter and Gilles Celeux a few years ago. I am therefore biased in what I would have included in a new book on the topic, the more because I find the available literature already plentiful, even though the early (1984) book of Titterington et al. that was my entry to the field may have become an historical reference. For instance, Finite Mixtures by McLachlan and Peel (2000) remains relevant, with similar emphasis on maximum likelihood and the EM algorithm, while Sylvia’s Finite Mixture and Markov Switching Models is still a reference to this day.

And an additional warning on me not being a massive fan of semi- and non-parametric estimation in this setting…

Preliminaries that may explain my limited enthusiasm about the book and its limited originality. Not that I found significant errors there (even though “improper priors [do not always] yield improper posteriors” [p.145] as we demonstrated in several papers), however, I had trouble with the uneven pace adopted by the authors that often skim some topics of importance while spending an inconsiderate amount of space on less relevant once. Some items get many bibliographical references, while others do not. For instance, EM receives a lion’s share (see, e..g, Sections 6.6 and 6.7). Or the 12 pages of proof in Chapter 10. Declination of sections into mixtures, mixtures of regressions, multivariate mixtures, hidden Markov models, and so on feels somewhat repetitive. This is particularly the case for the “mixture regression models” chapter.

The book also contains Bayesian entries, with a first introduction (p.105) in the discrete data chapter that precedes the short Bayesian chapter #4 (p.145), the same issue arising for related algorithms like Gibbs (p.107) that “estimate properties of the joint posterior” and MCMC (p.112). Which sort of erases the specificity of a Bayesian approach by reducing it to one item in the toolbox (with the wrong stress on MAP estimates). In this Bayesian chapter, MCMC validation is handled for discrete state spaces while applied in general spaces. The focus is mostly on relabelling for the following label switching chapter, albeit a large collection of methods are compared if not mentioned.

Handing an unknown number of components by hypothesis testing is supported in the next short chapter, although very little is said about reversible jump MCMC. And there is no general discussion on the consistency of these tests, in particular with bootstrap. Or at least on the regularity conditions they request. An puzzling paradox (p.191) is the existence of an unbounded Fisher information of an exponential mixture

\pi\mathcal Exp(1)+(1-\pi)\mathcal Exp(2)

when the weight π is the parameter (and close to 1).

High-dimensional mixtures in Chapter 8 are mostly handled by linear projections in smaller subspaces, which is natural given that they preserve the mixture structure but open a Pandora box of a wide range of proposed methods, again with little comparison available. Except in the R final section opposing several R functions on the same dataset (if unconclusively).

The semi-parametric chapters mention Dirichlet process priors, albeit briefly, but fail to relate to the recent works on using these when inferring about the number of components. Or failing to do so. There is also a very limited connection pointed out with machine learning but little can be gathered from the three page presentation (pp.308-310). These chapters also have significant overlap with the review paper of Xiang et al. (2019) in Statistical Science.

Most chapters end up with an R section, which usually reads as a quick demo of a related R package, like BayesLCA or our own mixtool. Hence not massively helpful beyond pointers to these packages. The numerical illustrations also are unevenly distributed between chapters, from nothing at all to four pages of small font tables on an MSE comparison between more or less robust approaches undertaken by Yu et al. (2020).

The above thus explains why I am not particularly excited about this bibliographical addition to the analysis of mixtures. It does offer a reference for researchers in the field by adding recent references and approaches to the existing books mentioned above, but I could not recommend it as a textbook (as suggested on p.xiii).

[Disclaimer about potential self-plagiarism: this post or an edited version may eventually appear in my Books Review section in CHANCE.]

identifying mixtures

Posted in Books, pictures, Statistics with tags , , , , , , on February 27, 2022 by xi'an

I had not read this 2017 discussion of Bayesian mixture estimation by Michael Betancourt before I found it mentioned in a recent paper. Where he re-explores the issue of identifiability and label switching in finite mixture models. Calling somewhat abusively degenerate mixtures where all components share the same family, e.g., mixtures of Gaussians. Illustrated by Stan code and output. This is rather traditional material, in that the non-identifiability of mixture components has been discussed in many papers and at least as many solutions proposed to overcome the difficulties of exploring the posterior distribution. Including our 2000 JASA paper with Gilles Celeux and Merrilee Hurn. With my favourite approach being the label-free representations as a point process in the parameter space (following an idea of Peter Green) or as a collection of clusters in the latent variable space. I am much less convinced by ordering constraints: while they formally differentiate and therefore identify the individual components of a mixture, they partition the parameter space with no regard towards the geometry of the posterior distribution. With in turn potential consequences on MCMC explorations of this fragmented surface that creates barriers for simulated Markov chains. Plus further difficulties with inferior but attracting modes in identifiable situations.

posterior collapse

Posted in Statistics with tags , , , , , , on February 24, 2022 by xi'an

The latest ABC One World webinar was a talk by Yixin Wang about the posterior collapse of auto-encoders, of which I was completely unaware. It is essentially an identifiability issue with auto-encoders, where the latent variable z at the source of the VAE does not impact the likelihood, assumed to be an exponential family with parameter depending on z and on θ, through possibly a neural network construct. The variational part comes from the parameter being estimated as θ⁰, via a variational approximation.

“….the problem of posterior collapse mainly arises from the model and the data, rather than from inference or optimization…”

The collapse means that the posterior for the latent satisfies p(z|θ⁰,x)=p(z), which is not a standard property since θ⁰=θ⁰(x). Which Yixin Wang, David Blei and John Cunningham show is equivalent to p(x|θ⁰,z)=p(x|θ⁰), i.e. z being unidentifiable. The above quote is then both correct and incorrect in that the choice of the inference approach, i.e. of the estimator θ⁰=θ⁰(x) has an impact on whether or not p(z|θ⁰,x)=p(z) holds. As acknowledged by the authors when describing “methods modify the optimization objectives or algorithms of VAE to avoid parameter values θ at which the latent variable is non-identifiable“. They later build a resolution for identifiable VAEs by imposing that the conditional p(x|θ,z) is injective in z for all values of θ. Resulting in a neural network with Brenier maps.

From a Bayesian perspective, I have difficulties to connect to the issue, the folk lore being that selecting a proper prior is a sufficient fix for avoiding non-identifiability, but more fundamentally I wonder at the relevance of inferring about the latent z’s and hence worrying about their identifiability or lack thereof.

One World ABC seminar [3.2.22]

Posted in Statistics, University life with tags , , , , , , , , , , , , on February 1, 2022 by xi'an

The next One World ABC seminar is on Thursday 03 Feb, with Yixing Want talking on Posterior collapse and latent variable non-identifiability It will take place at 15:30 CET (GMT+1).

Variational autoencoders model high-dimensional data by positing low-dimensional latent variables that are mapped through a flexible distribution parametrized by a neural network. Unfortunately, variational autoencoders often suffer from posterior collapse: the posterior of the latent variables is equal to its prior, rendering the variational autoencoder useless as a means to produce meaningful  epresentations. Existing approaches to posterior collapse often attribute it to the use of neural networks or optimization issues due to variational approximation. In this paper, we consider posterior collapse as a problem of latent variable non-identifiability. We prove that the posterior collapses if and only if the latent variables are non-identifiable in the generative model. This fact implies that posterior collapse is
not a phenomenon specific to the use of flexible distributions or approximate inference. Rather, it can occur in classical probabilistic models even with exact inference, which we also demonstrate. Based on these results, we propose a class of latent-identifiable variational autoencoders, deep generative models which enforce identifiability without sacrificing flexibility. This model class resolves the problem of latent variable non-identifiability by leveraging bijective Brenier maps and parameterizing them with input convex neural networks, without special variational inference objectives or optimization tricks. Across synthetic and real datasets, latent-identifiable variational  autoencoders outperform existing methods in mitigating posterior collapse and providing meaningful representations of the data.

multilevel linear models, Gibbs samplers, and multigrid decompositions

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , on October 22, 2021 by xi'an

A paper by Giacommo Zanella (formerly Warwick) and Gareth Roberts (Warwick) is about to appear in Bayesian Analysis and (still) open for discussion. It examines in great details the convergence properties of several Gibbs versions of the same hierarchical posterior for an ANOVA type linear model. Although this may sound like an old-timer opinion, I find it good to have Gibbs sampling back on track! And to have further attention to diagnose convergence! Also, even after all these years (!), it is always a surprise  for me to (re-)realise that different versions of Gibbs samplings may hugely differ in convergence properties.

At first, intuitively, I thought the options (1,0) (c) and (0,1) (d) should be similarly performing. But one is “more” hierarchical than the other. While the results exhibiting a theoretical ordering of these choices are impressive, I would suggest pursuing an random exploration of the various parameterisations in order to handle cases where an analytical ordering proves impossible. It would most likely produce a superior performance, as hinted at by Figure 4. (This alternative happens to be briefly mentioned in the Conclusion section.) The notion of choosing the optimal parameterisation at each step is indeed somewhat unrealistic in that the optimality zones exhibited in Figure 4 are unknown in a more general model than the Gaussian ANOVA model. Especially with a high number of parameters, parameterisations, and recombinations in the model (Section 7).

An idle question is about the extension to a more general hierarchical model where recentring is not feasible because of the non-linear nature of the parameters. Even though Gaussianity may not be such a restriction in that other exponential (if artificial) families keeping the ANOVA structure should work as well.

Theorem 1 is quite impressive and wide ranging. It also reminded (old) me of the interleaving properties and data augmentation versions of the early-day Gibbs. More to the point and to the current era, it offers more possibilities for coupling, parallelism, and increasing convergence. And for fighting dimension curses.

“in this context, imposing identifiability always improves the convergence properties of the Gibbs Sampler”

Another idle thought of mine is to wonder whether or not there is a limited number of reparameterisations. I think that by creating unidentifiable decompositions of (some) parameters, eg, μ=μ¹+μ²+.., one can unrestrictedly multiply the number of parameterisations. Instead of imposing hard identifiability constraints as in Section 4.2, my intuition was that this de-identification would increase the mixing behaviour but this somewhat clashes with the above (rigorous) statement from the authors. So I am proven wrong there!

Unless I missed something, I also wonder at different possible implementations of HMC depending on different parameterisations and whether or not the impact of parameterisation has been studied for HMC. (Which may be linked with Remark 2?)