Archive for cut distribution

BayesComp 2025.1

Posted in Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on June 18, 2025 by xi'an

Minus one day at BayesComp 2025! As I am attending the model misspecification satellite workshop (ten minutes late, due to repeated path finding protocol!), with an extended presentation by Jeremias Knoblauch on post-Bayesian inference, incl. powered likelihood and Gibbs posteriors. A very smooth and pedagogical presentation, esp. in the hybrid mode. A perspective I associate with the difficulties of making sense of the post-posterior, not truly a posterior, of calibrating the penalty (eg λ), picking the loss (α, β, γ divergences?) , and the drift towards learning goals since the new measure is the post-posterior predictive. Sort of paradoxical return to a Gaussian post-posterior on the parameter that does not seem to stay robust. Horrendous computational issues, when the loss itself is an integral. Use of the zig-zag sampler with an estimated unbiased gradient of the loss, much faster than pseudo-marginal, which (naïvely?) makes sense both because PDMPs directly use scores and because of the power of stochastic gradient methods. Worse perspectives for optimisation-centric posterior that are essentially vamped versions of GANs. For instance, what is the meaning of the coverage probabilities?

The second talk by Jonathan Huggins was on DC (not bagged) posteriors as martingale posteriors with m<∞ (approximating marginal distributions with random kernel MCMC—which persists in simulating the marginalised or integrated variable u from its prior, rather than adapting to the current value of the parameter θ— or subsampling MCMC akin to stochastic gradient Langevin) with connection with cut posteriors,

Then I skipped to the second workshop on Bayesian methods for distributional and semiparametric regression, to listen to my friend David Rossell’s talk on local variable selection. Which suffers more than in standard models under misspecification. Another talk involving cut posteriors, the cuts being on the spline bases…

The day and the workshop concluded with great talks by (my friends) Pierre Alquier and David Frazier. David centred his misspecification talk on cut posteriors. Managing to bring in shrinkage estimators (and mention Bill Strawderman!).

A wee stressful trip, since the races in Caen cancelled all buses and delayed the taxi enough to miss the train to Paris by 30s, catching the next available one leaving me less than one hour between the arrival of the train (delayed by construction work on the rail line) and boarding the flight at Charles de Gaulle airport, but fortunately the RER trains in Paris were running okay, there were no queues in the airport, and I thus made it in time with a bit of post-marathon jogging! (Only to be delayed at departure by one hour for stormy conditions over Germany and Austria). All this exercise proved helpful to sleep soundly and lengthily in the plane!

unbiased MCMC

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , on August 25, 2017 by xi'an

Two weeks ago, Pierre Jacob, John O’Leary, and Yves F. Atchadé arXived a paper on unbiased MCMC with coupling. Associating MCMC with unbiasedness is rather challenging since MCMC are rarely producing simulations from the exact target, unless specific tools like renewal can be produced in an efficient manner. (I supported the use of such renewal techniques as early as 1995, but later experiments led me to think renewal control was too rare an occurrence to consider it as a generic convergence assessment method.)

This new paper makes me think I had given up too easily! Here the central idea is coupling of two (MCMC) chains, associated with the debiasing formula used by Glynn and Rhee (2014) and already discussed here. Having the coupled chains meet at some time with probability one implies that the debiasing formula does not need a (random) stopping time. The coupling time is sufficient. Furthermore, several estimators can be derived from the same coupled Markov chain simulations, obtained by starting the averaging at a later time than the first iteration. The average of these (unbiased) averages results into a weighted estimate that weights more the later differences. Although coupling is also at the basis of perfect simulation methods, the analogy between this debiasing technique and perfect sampling is hard to fathom, since the coupling of two chains is not a perfect sampling instant. (Something obvious only in retrospect for me is that the variance of the resulting unbiased estimator is at best the variance of the original MCMC estimator.)

When discussing the implementation of coupling in Metropolis and Gibbs settings, the authors give a simple optimal coupling algorithm I was not aware of. Which is a form of accept-reject also found in perfect sampling I believe. (Renewal based on small sets makes an appearance on page 11.) I did not fully understood the way two random walk Metropolis steps are coupled, in that the normal proposals seem at odds with the boundedness constraints. But coupling is clearly working in this setting, while renewal does not. In toy examples like the (Efron and Morris!) baseball data and the (Gelfand and Smith!) pump failure data, the parameters k and m of the algorithm can be optimised against the variance of the averaged averages. And this approach comes highly useful in the case of the cut distribution,  a problem which I became aware of during MCMskiii and on which we are currently working with Pierre and others.