Following our arXival on the new version of our HPD based Gelfand & Dey estimator of evidence, I got pointed at Wang et al. (2018), which I had forgotten I had read at the time (as testified by an ‘Og entry). Reading my own comments, I concur (with myself¹⁸!) that the method is not massively compelling since it requires a partition set that is strongly related with the targeted integral. The above illustration for a mixture, that is for a pseudo posterior that is a mixture with two Gaussian components with known variance, also shows (in reverse) the curse of dimension and the need for finely tuned partitions. Said partition corresponding to the myriad of sets on the rhs. With such a degree of partitioning, Riemann integration should also produce perfect estimate, as shown by the zero error in the resulting estimator (Table 4).
Archive for curse of dimensionality
estimating evidence redux
Posted in Books, Statistics, University life with tags Bayesian Analysis, curse of dimensionality, estimating a constant, evidence, harmonic mean estimator, HPD region, importance sampling, marginal likelihood, Monte Carlo Statistical Methods on November 21, 2025 by xi'anApproximate Bayesian Computation with Statistical Distances for Model Selection [OWABI, 27 Nov]
Posted in Books, Statistics, University life with tags ABC model selection, Approximate Bayesian computation, approximate Bayesian inference, Bayesian inference, curse of dimensionality, information loss, intractable likelihood, One World Approximate Bayesian Inference Seminar, OWABI, simulation, simulation-based inference, summary statistics, toad, University of Warwick, webinar on November 17, 2025 by xi'an
The next OWABI seminar is delivered by Clara Grazian (University of Sidney), who will talk about “Approximate Bayesian Computation with Statistical Distances for Model Selection” on Thursday 27 November at 11am UK time:
Abstract: Model selection is a key task in statistics, playing a critical role across various scientific disciplines. While no model can fully capture the complexities of a real-world data-generating process, identifying the model that best approximates it can provide valuable insights. Bayesian statistics offers a flexible framework for model selection by updating prior beliefs as new data becomes available, allowing for ongoing refinement of candidate models. This is typically achieved by calculating posterior probabilities, which quantify the support for each model given the observed data. However, in cases where likelihood functions are intractable, exact computation of these posterior probabilities becomes infeasible. Approximate Bayesian computation (ABC) has emerged as a likelihood-free method and it is traditionally used with summary statistics to reduce data dimensionality, however this often results in information loss difficult to quantify, particularly in model selection contexts. Recent advancements propose the use of full data approaches based on statistical distances, offering a promising alternative that bypasses the need for handcrafted summary statistics and can yield posterior approximations that more closely reflect the true posterior under suitable conditions. Despite these developments, full data ABC approaches have not yet been widely applied to model selection problems. This paper seeks to address this gap by investigating the performance of ABC with statistical distances in model selection. Through simulation studies and an application to toad movement models, this work explores whether full data approaches can overcome the limitations of summary statistic-based ABC for model choice.
Keywords: model choice, distance metrics, full data approaches
Reference: C. Grazian, Approximate Bayesian Computation with Statistical Distances for Model Selection, preprint at ArXiv:2410.21603, 2025
Asymptotics of ABC when summaries converge at heterogeneous rates
Posted in pictures, Statistics, University life with tags ABC, Approximate Bayesian computation, Bayesian consistency, COVID-19, curse of dimensionality, lockdown, PhD thesis, summary statistics, Université Paris Dauphine, University of Oxford on November 21, 2023 by xi'an
We just posted a new arXival, jointly with Caroline Lawless, Judith Rousseau, and Robin Ryder. This is a significant component of Caroline’s PhD thesis in Oxford, on which we started working during the first COVID lockdown. In this paper, we extend our results with David Frazier, Gael Martin, both with whom I’ll soon be reunited!, and Judith, published in Biometrika in 2018, to the more challenging case where different components of the summary statistic vector converge to their respective means at different rates, with some possibly not even converging at all. While this sounds impossible (!), we do prove consistency of the ABC posterior under such heterogeneous rates.
Wentao Li and Paul Fearnhead (also in Biometrika and in 2018) reduce the curse of the dimension of the set of summary statistic by showing, in the specific case of asymptotically normal summary statistics concentrating at the same rate, that a local linear post-processing step leads to a significant improvement in the theoretical behaviour of the ABC posterior. However, due to this focus on reducing the impact of the dimension of the summary statistics, it is therefore important to study its efficiency in a context where the summary statistics are not as well behaved. Surprinsingly maybe, we show that the significant improvement due to local linear post-processing persists even when summary statistics have heterogeneous behaviour. Most interestingly, the number of summary statistics which converge at the fast rate has no impact on the rate of posterior concentration nor on the shape of the ABC posterior (provided it exceeds the dimension of the parameter).
prior elicitation
Posted in Books, Kids, Statistics, University life with tags ABC, Bayesian methods and expert elicitation, cognitive biases, conflicting prior, consensus prior, curse of dimensionality, prior elicitation, prior predictive, probabilistic programming, STAN, startup, summary statistics, whales, xkcd on January 13, 2022 by xi'an“We believe that an elicitation method should support elicitation both in the parameter and observable space, should be model-agnostic, and should be sample-efficient since human effort is costly.”
Petrus Mikkola et al. arXived a long paper on prior elicitation addressing the (most relevant) question: Why are we not widely use prior elicitation? With a massive bibliography that could be (partly) commented (and corrected as some references are incomplete, as eg my book chapter on priors!). I think the paper would make a terrific discussion paper.
The absence of a general procedure for prior elicitation is indeed hindering the adoption of Bayesian methods outside our core community and is thus eventually detrimental to their wider development. It also carries the dangers of misled or misleading prior choices. The authors put forward the absence of “software that integrates well with the current probabilistic programming tools used for other parts of the modelling workflow.” This requires setting principles that avoid “just-press-key” solutions. (Aside: This reminds me of my very first prospective PhD student, who was then working in a startup [although the name was not yet in use in the early 1990’s!] and had build such a software in a discretised, low dimension, conjugate prior, environment by returning a form of decision-theoretic impact of the chosen hyperparameters. He alas aborted his PhD attempt due to the short-term pressing matters in the under-staffed company…)
“We inspect prior elicitation from the perspectives of (1) properties of the prior distribution itself, (2) the model family and the prior elicitation method’s dependence on it, (3) the underlying elicitation space, (4) how the method interprets the information provided by the expert, (5) computation, (6) the form and quantity of interaction with the expert(s), and (7) the assumed capability of the expert (…)”
Prior elicitation is indeed a delicate balance between incorporating expert opinion(s) and avoiding over-standardisation. In my limited experience, experts tend to be over-confident about their own opinion and unwilling to attach uncertainty to their assessments. Even when being inconsistent. When several experts are involved (as, very briefly, in Section 3.6), building a common prior quickly becomes a challenge, esp. if their interests (or utility functions) diverge. As illustrated in the case of the whaling commission analysed by Adrian Raftery in the late 1990’s. (The above quote involves a single expert.) Actually, I dislike the term expert altogether, as it comes without any grading of the reliability of the person.
To hit (!) at an early statement in the paper (p.5), should the prior elicitation always depend on the (sampling) model, as experts may ignore or misapprehend the model? The posterior already accounts for the likelihood and the parameter may pre-exist wrt the model, as eg cosmological constants or vaccine efficiency… In a sense, the model should be involved as little as possible in the elicitation as the expert could confuse her beliefs about the parameter with those about the accuracy of the model. (I realise this is not necessarily a mainstream position as illustrated by this paper by Andrew and friends!)
And isn’t the first stumbling block the inability of most to represent one’s prior knowledge in probabilistic terms? Innumeracy is a shared shortcoming in the general population (and since everyone’s an expert!), as repeatedly demonstrated since the start of the Covid-19 pandemic. (See also the above point about inconsistency. Accounting for such inconsistencies in a Bayesian way is a natural answer, albeit requiring the degree of expertise and reliability to be tested.)
Is prior elicitation feasible beyond a few dimensions? Even when using the constrictive tool of copulas one hits a wall after a few dimensions, assuming the expert is willing to set a prior correlation matrix. Most of the methods described in Section 3.1 only apply to textbook examples. In their third dimension (!), the authors mention neural network parameters but later fail to cover this type of issue. (This was the example I had in mind indeed.) And they move from parameter space to observable space. Distinguishing predictive elicitation from observational elicitation, the former being what I would have suggested from scratch. Obviously, the curse of dimensionality strikes again unless one considers summary statistics (like in ABC).
While I am glad conjugate priors do not get the lion’s share, using as in Section 3.3.. non-parametric or machine learning solutions to construct the prior sounds unrealistic. (And including maximum entropy priors into that category seems wrong since they are definitely parametric.)
The proposed Bayesian treatment of the expert’s “data” (Section 4.1) is rational but requires an additional model construct to link the expert’s data with the parameter to reach a Bayes formula like (4.1). Plus a primary prior (which could then be one of the reference priors.) Reducing the expert’s input to imaginary observations may prove too narrow, though. The notion of an iterative elicitation is most appealing and its sequential aspect may not be particularly problematic in opposition to posteriors relying on using the data twice or more. I am much less buying the hierarchical construct of Section 4.3 because they imply a return to conjugate priors and hyperpriors, are not necessarily correctly understood by experts, do not always cater to observational elicitation, and are not an answer to high-dimension challenges.
Given the state of the art, it sounds like we are still far from seeing prior elicitation as a natural part of Bayesian software and probabilistic programming. Even when using a modular, model-agnostic strategy. But this is most certainly a worthy prospect!
improving bridge samplers by GANs
Posted in Books, pictures, Statistics with tags bridge sampling, curse of dimensionality, GANs, noise contrasting estimation, normalising flow, PhD, Saint Giles cemetery, University of Oxford on July 20, 2021 by xi'an
Hanwen Xing from Oxford recently posted a paper on arXiv about using GANs to improve the overlap bewtween the densities in bridge sampling. Bringing out new connections with noise contrastive estimation. The idea is to optimise a transform of one of the densities h() to bring it closer to the other density k(), using for instance normalising flows. (The call to transforms for bridge is not new, dating at least to Voter in 1985, the year I was starting my PhD!) Furthermore, using an f-divergence as a measure of functional distance allows for a reasonably straightforward update of the transform. That can be reformulated as a GAN target, which is somewhat natural in that the transform aims at confusing simulation from the transform of h and from k. This is quite an interesting proposal, even though calculating the optimal transform is time-consuming and subjet to the curse of dimensionality. I also wonder at whether or not iterating the optimisation, one density after the other, would be bring further improvement.