Archive for Markov chain

MaSeMo: Markov, Semi-Markov Models & Associated Fields [Paris, 1-4 July]

Posted in Statistics, University life with tags , , , , , , , , on May 7, 2025 by xi'an

integral priors for model comparison [2.0]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , on May 2, 2025 by xi'an

integral priors for multiple comparison

Posted in Books, Statistics, University life with tags , , , , , , , , , , , on June 24, 2024 by xi'an

Diego Salmerón and I just arXived a paper on integral priors for multiple model comparison, about deriving reference priors for multiple hypothesis testing. As (so-called) noninformative priors constructed for estimation purposes are usually not appropriate for model selection and testing due to their improperness, Jeffreys-Lindley paradoxes and the like, the methodology of integral priors was developed to get prior distributions for Bayesian model selection when comparing two models, modifying initial improper reference priors. This paper proposes a generalization of this methodology when than two models are to be compared. In order to avoid the above paradoxes and the associated possibility of producing a null recurrent or transient Markov chain, our approach adds an artificial copy of each model under comparison by compactifying the corresponding parametric space and creates an ergodic Markov chain exploring all models that returns the integral priors as marginals of the ergodic and stationary joint distribution. Besides the guarantee of existence of these integral priors and the disappearance of paradoxes that plague estimation reference priors, an additional perk of this methodology is that the simulation of this Markov chain is straightforward as it only requires simulations of imaginary training samples and from the corresponding posterior distributions, for all models, while producing Bayes factor approximations on the side. This renders its implementation automatic and generic, both in the nested and in the nonnested cases. We associated our late friend Juan Antonio Cano to this paper as he was instrumental in initiating both this collaboration and the methodology at its core.

joint fiddlin

Posted in Books, Kids, R, Statistics with tags , , , , , on April 22, 2024 by xi'an

Flip a fair coin 100 times, resulting in a sequence of heads (H) and tails (T). For each HH in the sequence, Alice gets a point; for each HT, Bob does, so e.g. for the subsequence THHHT Alice gets 2 points and Bob gets 1 point. Who is most likely to win?

An interesting conundrum in that the joint distribution of (A,B) need be considered for showing that Bob is more likely. Indeed, looking at the marginals does not help since the probability of the base events is the same. A solution on X validated (for a question posted when the Fiddler’s puzzle came out, Friday morn) demonstrates via a four state Markov chain representation the result (obvious from a quick simulation) that Alice wins 45% of the time while Bob wins 48%. The intuition is that, each time Alice wins at least a point, Bob gets an extra point at the end of the sequence (except possibly at the stopping time t=100), while in other cases Alice and Bob have the same probability to win one point.

null recurrent = zero utility?

Posted in Books, R, Statistics with tags , , , , , , , on April 28, 2022 by xi'an

The stability result that the ratio

\dfrac{\sum^T_{t=1} f(\theta^{(t)})}{\sum^T_{t=1} g(\theta^{(t)})}\qquad(1)

converges holds for a Harris π-null-recurrent Markov chain for all functions f,g in L¹(π) [Meyn & Tweedie, 1993, Theorem 17.3.2] is rather fascinating. However, it is unclear it can be useful in simulation environments, as for the integral priors we have been studying over the years with Juan Antonio Cano and Diego Salmeron Martinez. Above, the result of an experiment where I simulated a Markov chain as a Normal random walk in dimension one, hence a Harris π-null-recurrent Markov chain for the Lebesgue measure λ, and monitored the stabilisation of the ratio (1) when using two densities for f and g,  to its expected value (1, shown by a red horizontal line). There is quite a variability in the outcome (repeated 100 times),  but the most intriguing is the quick stabilisation of most cumulated averages to values different from 1. Even longer runs display this feature

which I would blame on the excursions of the random walk far away from the central regions for both f and g, that is on long sequences where zeroes keep being added to numerator and denominators in (1). As far as integral approximation is concerned, this is not very helpful!