amortized Bayesian mixture model

A few days before the January OWABI, I read through Simon Kucharsky’s and Paul Bürkner’s paper, arXived on 17 January. Which proposed an amortized Bayesian inference (ABI) method, even though the ABI is not the same as in OWABI! The motivation for their work is to start from a (standard) mixture model where the components are not analytically tractable (but still parameterised). But a generative model nonetheless. As in the earlier reviewed paper (which was arXived on the same day), by MEJ Newman, the dual representation of the joint posterior p(θ,z|x) as p(z|x,θ)p(θ|x) and p(θ|z,x)p(z|x) is (over?) emphasized (albeit unclearly why!). ABI uses neural networks and more specifically normalising flows to approximate the posterior p(θ|x) from prior predictive samples (θ,x) (as in ABC), and then directly exploit the invertibility of said flows to generate from this approximate posterior. One interesting aspect of the modelling is the derivation of summary statistics in the design of the network, albeit mixture posteriors do not allow for dimension-reduced (Bayes) sufficient statistics (and a contradictory sentence that conditioning on the summaries “does not alter the target posterior”, p7). The resulting approximate posterior generator proves much much faster than running an MCMC, obviously, and furthermore adapt to handling a sequence of datasets. A second network is constructed to approximate p(z|x,θ), using the same summaries. The network parameters are estimated through losses, rather than in a Bayesian manner, with a default Kullback-Leibler version (18). I also fail to understand why the networks are trained over unconstrained parameters when all parameters could become unconstrained when using the adequate parameterisation. And am fairly surprised at the regression towards the ill-fated step of using ordered parameters to avoid label switching… But the main quandary remains the issue of assessing the approximation effect, despite experiments aiming at pacifying such worries. And similarities with Stan and BayesFlow.

Leave a Reply

Discover more from Xi'an's Og

Subscribe now to keep reading and get access to the full archive.

Continue reading