Archive for Cornell University

miXtures on arXiv

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , on February 5, 2025 by xi'an

A paper about Bayesian inference on mixtures was posted on arXiv last week, as of 13 Jan 2025.  Fast sampling and model selection for Bayesian mixture models, by M. E. J. Newman is based on the notion that (genuine) parameters of a mixture model can be marginalized out when using conjugate priors. This is something that we pointed out quite a while ago, in a 1999 paper with George and Marty, which was devised in a long ride from Baltimore to Cornell after JSM 1999, and again in the 2002 Series B perfect sampling paper with George, Kerrie and Mike. (Also written in 1999.) And marginal likelihood can furthermore be approximated along this way as discussed in the more recent papers Bayesian Inference on Mixtures of Distributions with Kate, Kerrie & Jean-Michel, as well as Approximating the marginal likelihood in mixture models with Jean-Michel.

“Standard mixture models, as commonly formulated, also suffer from a technical, but important, difficulty: the existence of empty components. In many models (…) the number of observations in a component can be zero. Arguably this is acceptable for a model with a fixed number of components, but when the number of components is a free random variable it causes ambiguity, because a given division of observations into components can be represented in more than one way in the model. For instance, we could divide observations into two components, or we could divide them into three components, one of which is empty. This in turn creates difficulties when estimating the number of components—do we have two components or three?”

A very puzzling perspective, imho, since potentially empty components are inherent to (both finite and infinite) mixture models with connected issues of prohibiting some improper priors (if not all) and non-identifiability, including non-identifiability of the number of empty components (which remains random conditional on the data!), but different numbers of components lead to different models and their comparison is handled straightforwardly by a Bayesian analysis.

The author then proceeds to “prohibit empty components” [as a prior choice ?] as we did in the original (!) Gibbs sampler for mixtures in 1990 (published in 1994 in Series B!), seeking posterior properness, a trick later validated by Larry Wasserman (in again 1999, the year of mixtures!). Who called the construct the combination of a fixed prior and of a pseudo-likelihood, correctly imho (as the data dependent part is not properly normalised by a function of the parameters), rather than a prior choice. (The very one who stated that “mixtures, like tequila, are evil and should be avoided“.)

From there, the modelling is rather standard, with an arbitrary prior on k, number of components, a random partition model that prohibits empty components, even though the constraint could be more stringent depending on the number of parameters of a given component and the degree of improperness of the prior, as in our 1990 Series B paper. (Impropriety is not discussed in the paper.) Bayesian inference on k is based on the simulated (pseudo-)posterior. The choice therein as the estimated clustering is the most frequent partition (consensus clustering), connected to our proposal of (again!) 1999 with Merrilee and Gilles. While the estimated mixture is not explicited. The approach is assessed as running at an O(k) cost, with no parallel in terms of the data size n, even though the examples include a 59,946 dataset. One notable algorithmic trick when moving k is in selecting a component at random first rather than an observation index.

Some minor issues: detailed balance indicated as required for convergence (p14), label switching is called component switching (p5), higher acceptance rate indicated as meaning improved performances (p7)

support arXiv on Π

Posted in Books, University life with tags , , , , , on March 14, 2024 by xi'an

A. K. Md. Ehsanes Saleh (01 Jan 1932 – 03 Sept 2023)

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , , , , , on December 10, 2023 by xi'an

Just learned this day that Professor A. K. Md. Ehsanes Saleh passed away in early September. I first met him sometimes in the Fall of 1987, while visiting (from Purdue where I was visiting professor) my wife in Ottawa (where she was pursuing a Master in Electrical Engineering). I knew of his papers on shrinkage and pre-test estimators and dropped by Carleton University, where he taught and worked most of his life, for a casual talk. He was incredibly welcoming and friendly to an unknown junior researcher who had dropped by with no warning on a Friday afternoon. We then kept in touch about research projects and he made me an offer to visit Carleton over the Summer of 1988, with a welcome financial support that allowed us to rent a better lodging by the University of Ottawa (which my wife kept for the following year). This suited me most perfectly as I could spend the summer (May-August) with my wife and work with Professor Saleh on shrinkage topics, which was most enjoyable (if not immensely innovative), although the move involved a non-stop 14h drive from West Lafayette to Ottawa! The whole group of statisticians and probabilists at Carleton was unbelievably friendly as well and contributed, along with the stressless atmosphere of the Canadian capital and the endless nearby parks, to make that summer of 1988 a fabulous one. We renewed the experiment the following summer of 1989, when I left Cornell at the end of their semester, again a great one, when I also met Tatsuya Kubokawa who was visiting Professor Saleh as well. After those two years, I had very few opportunities to visit Ottawa and hence to meet him again, even though I remember having lunch with him at a Franco-Canadian meeting in 2008. I do and will remember him as a humble and selfless man, despite his accomplishments of being the first Bangladeshi statistician in receiving many awards and distinctions, always amicable and full of tolerance and helpful advice.

Hubert Reeves (1932-2023)

Posted in Books, Kids, pictures, University life with tags , , , , , , , on October 14, 2023 by xi'an

arX[g]iv[e]

Posted in Statistics with tags , on March 16, 2023 by xi'an