Archive for ergodicity
on(-line) integral priors for model selection
Posted in Books, Statistics, University life with tags Bayes factor, Bayesian model selection, collaboration, ergodicity, improper priors, integral priors, International Statistical Review, ISI, Juan Antonio Cano, Markov chains, MCMC, noninformative priors, open access, paper, reference priors on February 27, 2026 by xi'anintegral priors for model comparison [2.0]
Posted in Books, Statistics, University life with tags arXiv, Bayesian inference, Bayesian model choice, Bayesian statistics, ergodicity, imaginary training sample, integral priors, Juan Antonio Cano, Markov chain, MCMC, Murcia, nested models, noninformative priors, reference priors, Spain, Universidad de Murcia on May 2, 2025 by xi'anintegral priors for multiple comparison
Posted in Books, Statistics, University life with tags ergodicity, imaginary training sample, improper posteriors, improper priors, intrinsic Bayes factor, Juan Antonio Cano, Markov chain, noninformative priors, null recurrence, pseudo-Bayes factors, reference priors, transience on June 24, 2024 by xi'an
Diego Salmerón and I just arXived a paper on integral priors for multiple model comparison, about deriving reference priors for multiple hypothesis testing. As (so-called) noninformative priors constructed for estimation purposes are usually not appropriate for model selection and testing due to their improperness, Jeffreys-Lindley paradoxes and the like, the methodology of integral priors was developed to get prior distributions for Bayesian model selection when comparing two models, modifying initial improper reference priors. This paper proposes a generalization of this methodology when than two models are to be compared. In order to avoid the above paradoxes and the associated possibility of producing a null recurrent or transient Markov chain, our approach adds an artificial copy of each model under comparison by compactifying the corresponding parametric space and creates an ergodic Markov chain exploring all models that returns the integral priors as marginals of the ergodic and stationary joint distribution. Besides the guarantee of existence of these integral priors and the disappearance of paradoxes that plague estimation reference priors, an additional perk of this methodology is that the simulation of this Markov chain is straightforward as it only requires simulations of imaginary training samples and from the corresponding posterior distributions, for all models, while producing Bayes factor approximations on the side. This renders its implementation automatic and generic, both in the nested and in the nonnested cases. We associated our late friend Juan Antonio Cano to this paper as he was instrumental in initiating both this collaboration and the methodology at its core.
exact yet private MCMC
Posted in Statistics with tags Arrowleaf Cellars, differential privacy, ergodicity, ICML 2023, Lake Okanagan, MCMC, Metropolis-Hastings algorithm, Okanagan vineyards, Poisson subsampling, reversibility, spectral gap, stationarity on August 9, 2023 by xi'an
“at each iteration, DP-fast MH first samples a minibatch size and checks if it uses a minibatch of data or full-batch data. Then it checks whether to require additional Gaussian noise. If so, it will instantiate the Gaussian mechanism which adds Gaussian noise to the energy difference function. Finally, it chooses accept or reject θ′ based on the noisy acceptance probability.”
Private, Fast, and Accurate Metropolis-Hastings for Large-Scale Bayesian Inference is an(other) ICML²³ paper, written by Wanrong Zhang and Ruqi Zhang. Who are running MCMC under DP constraints. For one thing, they compute the MH acceptance probability with a minibatch, which is Poisson sampled (in order to guarantee privacy). It appears as a highly calibrated algorithm (see, e.g., Algorithm 1). Under the assumption (1) that the difference between individual log densities for two values of the parameter is upper bounded (in the data), differential privacy is established as failing to detect for certain a datapoint from the MCMC output. Interestingly, the usual randomisation leading to pricacy is operated on the energy level, rather than on observations or summary statistics, although this may prove superfluous when there is enough randomness provided by the MH step itself: “inherent privacy guarantees in the MH algorithm”
“when either the privacy hyperparameter ϵ or δ becomes small, the convergence rate becomes small, characterizing how much the privacy constraint slows down the convergence speed of the Markov chain”
The major results of the paper are privacy guarantees (at each iteration) and preservation of the proper target distribution, in contrast with earlier versions. In particular, adding the Gaussian noise to the energy does not impact reversibility. (Even though I am not 100% sure I buy the entire argument about reversibility (in Appendix C) as it sounds too easy!) The authors even achieve a bound on the relative spectral gaps.
Bayes Rules! [book review]
Posted in Books, Kids, Mountains, pictures, R, Running, Statistics, University life with tags Abele Blanc, Adriana Rosenbluth, Ama Dablam, Annapurna, applied Bayesian analysis, Australia, Bayes factor, Bayes rule, Bayesian modelling, Bletchley Park, book ban, book review, CHANCE, conjugate priors, CRAN, effective sample size, Enigma code machine, ergodicity, Florida, Gibbs distribution, Himalayas, history of statistics, introductory textbooks, it's greek to me, MCMC, Monte Carlo Statistical Methods, R, rStan, simulation, STAN, Uluru, undergraduates, weakly informative prior on July 5, 2022 by xi'an
Bayes Rules! is a new introductory textbook on Applied Bayesian Model(l)ing, written by Alicia Johnson (Macalester College), Miles Ott (Johnson & Johnson), and Mine Dogucu (University of California Irvine). Textbook sent to me by CRC Press for review. It is available (free) online as a website and has a github site, as well as a bayesrule R package. (Which reminds me that both our own book R packages, bayess and mcsm, have gone obsolete on CRAN! And that I should find time to figure out the issue for an upgrading…)
As far as I can tell [from abroad and from only teaching students with a math background], Bayes Rules! seems to be catering to early (US) undergraduate students with very little exposure to mathematical statistics or probability, as it introduces basic probability notions like pmf, joint distribution, and Bayes’ theorem (as well as Greek letters!) and shies away from integration or algebra (a covariance matrix occurs on page 437 with a lot . For instance, the Normal-Normal conjugacy derivation is considered a “mouthful” (page 113). The exposition is somewhat stretched along the 500⁺ pages as a result, imho, which is presumably a feature shared with most textbooks at this level, and, accordingly, the exercises and quizzes are more about intuition and reproducing the contents of the chapter than technical. In fact, I did not spot there a mention of sufficiency, consistency, posterior concentration (almost made on page 113), improper priors, ergodicity, irreducibility, &tc., while other notions are not precisely defined, like ESS, weakly informative (page 234) or vague priors (page 77), prior information—which makes the negative answer to the quiz “All priors are informative” (page 90) rather confusing—, R-hat, density plot, scaled likelihood, and more.
As an alternative to “technical derivations” Bayes Rules! centres on intuition and simulation (yay!) via its bayesrule R package. Itself relying on rstan. Learning from example (as R code is always provided), the book proceeds through conjugate priors, MCMC (Metropolis-Hasting) methods, regression models, and hierarchical regression models. Quite impressive given the limited prerequisites set by the authors. (I appreciated the representations of the prior-likelihood-posterior, especially in the sequential case.)
Regarding the “hot tip” (page 108) that the posterior mean always stands between the prior mean and the data mean, this should be made conditional on a conjugate setting and a mean parameterisation. Defining MCMC as a method that produces a sequence of realisations that are not from the target makes a point, except of course that there are settings where the realisations are from the target, for instance after a renewal event. Tuning MCMC should remain a partial mystery to readers after reading Chapter 7 as the Goldilocks principle is quite vague. Similarly, the derivation of the hyperparameters in a novel setting (not covered by the book) should prove a challenge, even though the readers are encouraged to “go forth and do some Bayes things” (page 509).
While Bayes factors are supported for some hypothesis testing (with no point null), model comparison follows more exploratory methods like X validation and expected log-predictive comparison.
The examples and exercises are diverse (if mostly US centric), modern (including cultural references that completely escape me), and often reflect on the authors’ societal concerns. In particular, their concern about a fair use of the inferred models is preminent, even though a quantitative assessment of the degree of fairness would require a much more advanced perspective than the book allows… (In that respect, Exercise 18.2 and the following ones are about book banning (in the US). Given the progressive tone of the book, and the recent ban of math textbooks in the US, I wonder if some conservative boards would consider banning it!) Concerning the Himalaya submitting running example (Chapters 18 & 19), where the probability to summit is conditional on the age of the climber and the use of additional oxygen, I am somewhat surprised that the altitude of the targeted peak is not included as a covariate. For instance, Ama Dablam (6848 m) is compared with Annapurna I (8091 m), which has the highest fatality-to-summit ratio (38%) of all. This should matter more than age: the Aosta guide Abele Blanc climbed Annapurna without oxygen at age 57! More to the point, the (practical) detailed examples do not bring unexpected conclusions, as for instance the fact that runners [thrice alas!] tend to slow down with age.
A geographical comment: Uluru (page 267) is not a city!, but an impressive sandstone monolith in the heart of Australia, a 5 hours drive away from Alice Springs. And historical mentions: Alan Turing (page 10) and the team at Bletchley Park indeed used Bayes factors (and sequential analysis) in cracking the Enigma, but this remained classified information for quite a while. Arianna Rosenbluth (page 10, but missing on page 165) was indeed a major contributor to Metropolis et al. (1953, not cited), but would not qualify as a Bayesian statistician as the goal of their algorithm was a characterisation of the Boltzman (or Gibbs) distribution, not statistical inference. And David Blackwell’s (page 10) Basic Statistics is possibly the earliest instance of an introductory Bayesian and decision-theory textbook, but it never mentions Bayes or Bayesianism.
[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Book Review section in CHANCE.]


“at each iteration, DP-fast MH first samples a minibatch size and checks if it uses a minibatch of data or full-batch data. Then it checks whether to require additional Gaussian noise. If so, it will instantiate the Gaussian mechanism which adds Gaussian noise to the energy difference function. Finally, it chooses accept or reject θ′ based on the noisy acceptance probability.”