Archive for hidden Markov models

mixture models [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , on August 14, 2024 by xi'an

Strangely enough, I became aware of this new book on mixtures through one of these annoying emails “Your work has been cited n times this week“… Mixture Models (Parametric, Semiparametric, and New Directions) by Weixin Yao and Sijia Wang got published by CRC Press earlier this year, within the Monographs on Statistics and Applied Probability green series (#175), and covers across 380 pages most aspects of mixture (and hidden Markov) estimation, if with strong emphasis on maximum likelihood estimation, while the new directions are unsurprisingly those pursued by the authors, namely robust and semi-parametric estimation, as well as model selection by testing.

An early warning about this book review is that I co-edited a Handbook of Mixture Analysis with my friends Sylvia Früwirth-Schnatter and Gilles Celeux a few years ago. I am therefore biased in what I would have included in a new book on the topic, the more because I find the available literature already plentiful, even though the early (1984) book of Titterington et al. that was my entry to the field may have become an historical reference. For instance, Finite Mixtures by McLachlan and Peel (2000) remains relevant, with similar emphasis on maximum likelihood and the EM algorithm, while Sylvia’s Finite Mixture and Markov Switching Models is still a reference to this day.

And an additional warning on me not being a massive fan of semi- and non-parametric estimation in this setting…

Preliminaries that may explain my limited enthusiasm about the book and its limited originality. Not that I found significant errors there (even though “improper priors [do not always] yield improper posteriors” [p.145] as we demonstrated in several papers), however, I had trouble with the uneven pace adopted by the authors that often skim some topics of importance while spending an inconsiderate amount of space on less relevant once. Some items get many bibliographical references, while others do not. For instance, EM receives a lion’s share (see, e..g, Sections 6.6 and 6.7). Or the 12 pages of proof in Chapter 10. Declination of sections into mixtures, mixtures of regressions, multivariate mixtures, hidden Markov models, and so on feels somewhat repetitive. This is particularly the case for the “mixture regression models” chapter.

The book also contains Bayesian entries, with a first introduction (p.105) in the discrete data chapter that precedes the short Bayesian chapter #4 (p.145), the same issue arising for related algorithms like Gibbs (p.107) that “estimate properties of the joint posterior” and MCMC (p.112). Which sort of erases the specificity of a Bayesian approach by reducing it to one item in the toolbox (with the wrong stress on MAP estimates). In this Bayesian chapter, MCMC validation is handled for discrete state spaces while applied in general spaces. The focus is mostly on relabelling for the following label switching chapter, albeit a large collection of methods are compared if not mentioned.

Handing an unknown number of components by hypothesis testing is supported in the next short chapter, although very little is said about reversible jump MCMC. And there is no general discussion on the consistency of these tests, in particular with bootstrap. Or at least on the regularity conditions they request. An puzzling paradox (p.191) is the existence of an unbounded Fisher information of an exponential mixture

\pi\mathcal Exp(1)+(1-\pi)\mathcal Exp(2)

when the weight π is the parameter (and close to 1).

High-dimensional mixtures in Chapter 8 are mostly handled by linear projections in smaller subspaces, which is natural given that they preserve the mixture structure but open a Pandora box of a wide range of proposed methods, again with little comparison available. Except in the R final section opposing several R functions on the same dataset (if unconclusively).

The semi-parametric chapters mention Dirichlet process priors, albeit briefly, but fail to relate to the recent works on using these when inferring about the number of components. Or failing to do so. There is also a very limited connection pointed out with machine learning but little can be gathered from the three page presentation (pp.308-310). These chapters also have significant overlap with the review paper of Xiang et al. (2019) in Statistical Science.

Most chapters end up with an R section, which usually reads as a quick demo of a related R package, like BayesLCA or our own mixtool. Hence not massively helpful beyond pointers to these packages. The numerical illustrations also are unevenly distributed between chapters, from nothing at all to four pages of small font tables on an MSE comparison between more or less robust approaches undertaken by Yu et al. (2020).

The above thus explains why I am not particularly excited about this bibliographical addition to the analysis of mixtures. It does offer a reference for researchers in the field by adding recent references and approaches to the existing books mentioned above, but I could not recommend it as a textbook (as suggested on p.xiii).

[Disclaimer about potential self-plagiarism: this post or an edited version may eventually appear in my Books Review section in CHANCE.]

All About that Bayes stroll

Posted in pictures, Statistics, University life with tags , , , , , , , , , , , , , on February 9, 2024 by xi'an

For all Bayesians and sympathisers in the Paris area, an incoming All about that Bayes seminars¹ by Elisabeth Gassiat (Institut de Mathématiques d’Orsay) on 13 February, 16h00, on Campus Pierre & Marie Curie, SCAI:

A stroll through hidden Markov models

Hidden Markov models are latent variables models producing dependent sequences. I will survey recent results providing guarantees for their use in various fields such as clustering, multiple testing, nonlinear ICA or variational autoencoders.


¹Incidentally, I came across an unrelated All about that Bayes YouTube video, a talk given by Kristin Lennox (Lawrence Livermore National Laboratory). And then found out a myriad of talks or courses using that pun.

simulation based composite likelihood

Posted in Statistics with tags , , , , , on December 29, 2023 by xi'an

Lorenzo Rimella, Chris Jewell, and Paul Fearnhead have recently arXived a paper entitled Simulation Based Composite Likelihood, where they consider a composite likelihood approximation for running inference on HMM parameters under the specific scenario of HMMs on finite, high-dimension N, state spaces X with huge cost of order card (Χ)2N when computing the likelihood  by the forward algorithm:

“Inference for high-dimensional hidden Markov models is challenging due to the exponential-in-dimension computational cost of the forward algorithm.”

The authors make an assumption (2) of total factorisation across dimensions for both current hidden and current observed terms, given the previous hidden states, which is very very strong, if not resulting in a complete separation into independent component-wise HMMs. This helps however in deriving a Monte Carlo approximation of the likelihood of one component of the HMM sequence, the full likelihood being then approximated in a composite (likelihood) manner by the product of these component marginals.  The remaining difficulty of computing the marginals of the component-wise observed (pseudo-) Markov chains is attenuated

“by fixing the state of all but one component n of the latent process, [since] we can leverage the factorisation and calculate probabilities related to the time-trajectory of the remaining [latent] state”

but it requires simulation of the hidden chain, overall of order  O(PTN²card (X)²) when P is the number of MCMC simulations, which can be improved by a factor N by removing a feedback step through a further marginal likelihood approximation. Interestingly falling into a prediction-correction pattern usual in sequential simulations. All this demonstrates craftsmanship of a high order, even though the issue of using an approximate composite likelihood does not seem to be addressed.

 

Natural statistical science [#2]

Posted in Statistics with tags , , , , , , , , , , , , , , , on November 23, 2023 by xi'an

A rare occurrence of a Bayesian statistics paper in Nature with this “State estimation of a physical system with unknown governing equations” by Course and Nair. A variational Bayes modelling of a state system observed with noise, but without a physical model on the state (SDE) evolution itself. Which means a prior is set on a non-parametric or neural representation of the drift and a linear approximation is used for the variational approximation, leading to a Gaussian process as the approximate distribution. While this applies to highly complex models, like orbiting black holes, it is somewhat a surprise to meet this application of variational inference in a prestigious general science journal like Nature. (The picture above was taken on the train from Marseille at the end of the Bayes Fall school.)

“The approach is based on a technique called Bayesian inference, which is used widely, but which can be computationally challenging for complex systems.” B. Keith

Approximation Methods in Bayesian Analysis [#2]

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , on June 22, 2023 by xi'an

A more theoretical Day #2 of the workshop, with Debdeep Pati comparing two representations of Gaussian processes with significantly different efficiencies, and Aad van der Vaart presenting a form of linearisation for a range of inverse problems, Kolyan Ray debiasing Lasso impacts by variational Bayes, although through a somewhat intricate process that distanced the procedure from Bayesian grounds imho, Judith Rousseau (Dauphine) also drifting away from Bayesian canons by looking anew at empirical Bayes with surprising differences from genuine B analysis, connecting with the cutoff phenomenon she and Kerrie exhibited in their 2011 mixture paper, as well as labelling the marginal likelihood a misspecified model. Trevor Campbell and Sinead Williamson both provided Bayesian perspectives on normalising flows, in particular the impact of computer imprecision on reversibility, leading to the notion of shadow paths (screenshot below), while Giovanni Rebaudo talked about mixtures supported by trees, a fascinating object!

On Day #3, Marc Beaumont talked on a mixture of composite likelihood à la Ryden, making me wonder of optimisation of blocks for HMC? EP-ABC, with the issue of the unknown amount of approximation, and adaptivity?, Maria de Iorio presented work on finite and infinite mixtures with repulsive (Coulomb) priors, achieving a unified framework, plus known evidence (?), with a correlated talk by Federico Camerlenghi in the afternoon, with novel notions (for me) of Palm measures and calculus, and another correlated talk by María-Fernanda Gil Leyva Villa, on stick-breaking processes for species sampling with dependent length variables, with related improvements in Gibbs implementation (screenshot below).
This was followed by two theoretical talks on continuous time processes by Paul Jenkins (Warwick) on the fine properties of the Flemming-Viot process, with mentions of Don Dawson’s results reminding me of the 1988 and 1989 summers I spent at Carleton University, where he was located at the time, and Matteo Ruggieri, with the novel (to me) notion of dual Markov processes that could prove useful in a lot of latent variable models. Fabrizio Leisen expanded on his early work on partial exchangeability and Steve MacEachern on dependent quantile pyramids, which relate to quantile regression, a constant source of puzzlement for me. Motivating the perspective by robustness and misspecification arguments. But I am a wee bit puzzled by the distinction between quantile pyramids and other non-parametric solutions.

On the outdoor front (in early mornings), choppy waters at sea (in the Sugiton calanque, pictured above) thanks to the endless mistral wind, nice run down from Mont Puget with friends, limited utility of my rented mountain bike (except to reach the nearest supermarket, 3km away)