parallelising MCMC via random forests

tempWe have just arXived a new paper written with my former PhD student Wu Changye on the use of random forest regressions to learn the partial posteriors simulated in divide-and-conquer MCMC, when the whole data set into batches, runs MCMC algorithms separately over each batch to produce samples of parameters. Here, we use each resulting subposterior as a proposal distribution to implement importance sampling. Unlike the existing divide-and-conquer MCMC, our methods are based on scaled subposteriors, whose scale factors are not necessarily restricted to being equal to one or to the number of subsets. Given suitable scale factors, we can achieve overlapping subposteriors, a fea-ture that is of the highest importance in the combination stage. While we have no theoretical argument to promote this approach over existing divide-and-conquer MCMC techniques, it behaves satisfactorily over benchmark models, including (some) misspecified models. The main limitation of this approach is a curse of dimensionality at the random forest training stage. For higher dimensional parameter spaces, it does require more sample points to train each random forest learner.

Leave a Reply

Discover more from Xi'an's Og

Subscribe now to keep reading and get access to the full archive.

Continue reading