Archive for doubly intractable posterior

amortized Bayesian mixture model

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , on February 7, 2025 by xi'an

A few days before the January OWABI, I read through Simon Kucharsky’s and Paul Bürkner’s paper, arXived on 17 January. Which proposed an amortized Bayesian inference (ABI) method, even though the ABI is not the same as in OWABI! The motivation for their work is to start from a (standard) mixture model where the components are not analytically tractable (but still parameterised). But a generative model nonetheless. As in the earlier reviewed paper (which was arXived on the same day), by MEJ Newman, the dual representation of the joint posterior p(θ,z|x) as p(z|x,θ)p(θ|x) and p(θ|z,x)p(z|x) is (over?) emphasized (albeit unclearly why!). ABI uses neural networks and more specifically normalising flows to approximate the posterior p(θ|x) from prior predictive samples (θ,x) (as in ABC), and then directly exploit the invertibility of said flows to generate from this approximate posterior. One interesting aspect of the modelling is the derivation of summary statistics in the design of the network, albeit mixture posteriors do not allow for dimension-reduced (Bayes) sufficient statistics (and a contradictory sentence that conditioning on the summaries “does not alter the target posterior”, p7). The resulting approximate posterior generator proves much much faster than running an MCMC, obviously, and furthermore adapt to handling a sequence of datasets. A second network is constructed to approximate p(z|x,θ), using the same summaries. The network parameters are estimated through losses, rather than in a Bayesian manner, with a default Kullback-Leibler version (18). I also fail to understand why the networks are trained over unconstrained parameters when all parameters could become unconstrained when using the adequate parameterisation. And am fairly surprised at the regression towards the ill-fated step of using ordered parameters to avoid label switching… But the main quandary remains the issue of assessing the approximation effect, despite experiments aiming at pacifying such worries. And similarities with Stan and BayesFlow.

JSM 2024, Portland, Day 2

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , on August 7, 2024 by xi'an

By happenstance, I started my day in the cybersecurity session. With (again) hardly a soul in the room… A first talk on avoiding herding and achieving asymptotic truth learning (about a binary outcome) in a graph by putting constraints on the graph structure, without any clear connection with statistics or cybersecurity. Even less for the second talk on optimising masks. Only with the third one came cybersecurity motivations, the focus being on a two-player Stackelberg game already used in this framework. The result proper was about estimating the parameter of a (rather unrealistic) posited model reproducing the adversarial actions. The last talk about jailbreak attacks was again off-field by miles.


Then attended (as intended!) the 2024 Blackwell-Rosenbluth Award session, featuring the nominees Sharmistha Guha (Texas A&M) on multiple network inference, Simon Mak (Duke) on using Bayesian surrogate models, Guanyang Wang (Rutgers) who recently spoke at our mostly Monte Carlo seminar, Akihiko Nishimura (John Hopkins) on a unification of HMC and PDMPs, most appropriate when located next to Mount Hamilton and the Zigzag river!, and Maria Skoularidou (MIT) on evolutionary Monte Carlo for gene expression. Alas with hardly anyone in the room.

In the afternoon, I went to the (very well-attended this time!, with no seat available for many attendees, incl. yours truly!) COPSS Elizabeth L. Scott Lecture by my friend from Rutgers, Regina Liu, on the highly relevant challenge of combining inferences from diverse data sources. Using the (definitely Rutgerian!) approach of confidence distributions!

Last (late) afternoon, I went swimming from Kevin Duckworth dock, just below the conference centre. Water was quite warm (and green), with a few other swimmers, and no stomachical after-effect, so far. Hence I returned there once again this afternoon.

re-MCM’d

Posted in Statistics, University life with tags , , , , , , , , , , , , , , , on June 28, 2023 by xi'an


When I entered the classroom on the Jussieu campus where the Monday morning session on PDMP was taking place, some friends told me a badge was waiting for me at the registration desk of MCM 2023! Most surprisingly since I had received a deregistration message a few weeks earlier. Great PDMP session with non-reversible tempering (warning: some PDMPs were harmed in the tempering process), scaling PDMPs for highly anisotropic targets, a PDMP form of reversible jump (with sticky floors!) and a more adaptive version of no-U turn via… PDMPs. After lunch at the nearby Grande Mosquée de Paris (I had not visited since… 1974!), I attended the slice sampler session, where mileage varied imho. With a take on doubly-intractable targets that did not seem to relate to the existing literature.

scalable Metropolis-Hastings, nested Monte Carlo, and normalising flows

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , on June 16, 2020 by xi'an

Over a sunny if quarantined Sunday, I started reading the PhD dissertation of Rob Cornish, Oxford University, as I am the external member of his viva committee. Ending up in a highly pleasant afternoon discussing this thesis over a (remote) viva yesterday. (If bemoaning a lost opportunity to visit Oxford!) The introduction to the viva was most helpful and set the results within the different time and geographical zones of the Ph.D since Rob had to switch from one group of advisors in Engineering to another group in Statistics. Plus an encompassing prospective discussion, expressing pessimism at exact MCMC for complex models and looking forward further advances in probabilistic programming.

Made of three papers, the thesis includes this ICML 2019 [remember the era when there were conferences?!] paper on scalable Metropolis-Hastings, by Rob Cornish, Paul Vanetti, Alexandre Bouchard-Côté, Georges Deligiannidis, and Arnaud Doucet, which I commented last year. Which achieves a remarkable and paradoxical O(1/√n) cost per iteration, provided (global) lower bounds are found on the (local) Metropolis-Hastings acceptance probabilities since they allow for Poisson thinning à la Devroye (1986) and  second order Taylor expansions constructed for all components of the target, with the third order derivatives providing bounds. However, the variability of the acceptance probability gets higher, which induces a longer but still manageable if the concentration of the posterior is in tune with the Bernstein von Mises asymptotics. I had not paid enough attention in my first read at the strong theoretical justification for the method, relying on the convergence of MAP estimates in well- and (some) mis-specified settings. Now, I would have liked to see the paper dealing with a more complex problem that logistic regression.

The second paper in the thesis is an ICML 2018 proceeding by Tom Rainforth, Robert Cornish, Hongseok Yang, Andrew Warrington, and Frank Wood, which considers Monte Carlo problems involving several nested expectations in a non-linear manner, meaning that (a) several levels of Monte Carlo approximations are required, with associated asymptotics, and (b) the resulting overall estimator is biased. This includes common doubly intractable posteriors, obviously, as well as (Bayesian) design and control problems. [And it has nothing to do with nested sampling.] The resolution chosen by the authors is strictly plug-in, in that they replace each level in the nesting with a Monte Carlo substitute and do not attempt to reduce the bias. Which means a wide range of solutions (other than the plug-in one) could have been investigated, including bootstrap maybe. For instance, Bayesian design is presented as an application of the approach, but since it relies on the log-evidence, there exist several versions for estimating (unbiasedly) this log-evidence. Similarly, the Forsythe-von Neumann technique applies to arbitrary transforms of a primary integral. The central discussion dwells on the optimal choice of the volume of simulations at each level, optimal in terms of asymptotic MSE. Or rather asymptotic bound on the MSE. The interesting result being that the outer expectation requires the square of the number of simulations for the other expectations. Which all need converge to infinity. A trick in finding an estimator for a polynomial transform reminded me of the SAME algorithm in that it duplicated the simulations as many times as the highest power of the polynomial. (The ‘Og briefly reported on this paper… four years ago.)

The third and last part of the thesis is a proposal [to appear in ICML 20] on relaxing bijectivity constraints in normalising flows with continuously index flows. (Or CIF. As Rob made a joke about this cleaning brand, let me add (?) to that joke by mentioning that looking at CIF and bijections is less dangerous in a Trump cum COVID era at CIF and injections!) With Anthony Caterini, George Deligiannidis and Arnaud Doucet as co-authors. I am much less familiar with this area and hence a wee bit puzzled at the purpose of removing what I understand to be an appealing side of normalising flows, namely to produce a manageable representation of density functions as a combination of bijective and differentiable functions of a baseline random vector, like a standard Normal vector. The argument made in the paper is that imposing this representation of the density imposes a constraint on the topology of its support since said support is homeomorphic to the support of the baseline random vector. While the supporting theoretical argument is a mathematical theorem that shows the Lipschitz bound on the transform should be infinity in the case the supports are topologically different, these arguments may be overly theoretical when faced with the practical implications of the replacement strategy. I somewhat miss its overall strength given that the whole point seems to be in approximating a density function, based on a finite sample.