Archive for ABC variable selection

Mark [a]B[c], plus cats

Posted in Books, Kids, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , on June 22, 2024 by xi'an

30 May was a day of first and last times, if not in capital ways (for me), on the One World ABC webinar. This was the first time we had a talk by Mark Beaumont and also the first time I had team experience of facing a smoking participant, while this was the last time of our monthly webinar for the (Northern) academic year.

The talk was about model misspecification in population genomic, from an ABC perspective with the motivation of common noticeable difference between the distributions of d(s,s⁰) and d(s,s’), distances between the prior predictively simulated summary statistics and the observed ones vs posterior generated ones, which should indicates misspecification, esp with complicated models. Mark and his coauthors then supported a gradual elimination of summary statistics to diminish the discrepancy, hence voluntarily impoverishing the model. While blaming the statistics sounded a bit like shooting the messenger, the resolution is of obvious interest if backing from modelling the misspecification itself.

Some of the presented work was conducted in Ward et al (2022, NeurIPS) with a reference to the outlying Ratmann et al (2009) we later discussed, for including tolerance as an extra parameter ε, thereby reconsidering Wikinson’s exact ABC for noisy observations y by the medium of a normalising flow on the marginal distribution of the denoised x (learned from the prior predictive)

More precisely, the idea is to drop summary statistics by checking whether or not the observed S⁰ belongs to HPD region, removing one component of S at a time, using e.g. a k-NN estimate for the summary density (hence depending on parameterisation of said statistics for the distance). Hopefully, the process stops before loosing identifiability by using too few statistics. I also wondered at multiple uses of the data in this sequential procedure but Mark argued for adopting a meta- or pragma- or Gelmanian- Bayesian perspective in the end!

Another perk was the appearance of (and illustration with) the Scottish Wildcat, mentioned in The Guardian a few months ago and discussed in the ‘Og, with further papers exploring more aspects of this hybridization, like a posterior applied to a much more complex phylogenic tree reconstruction for cats of different creeds and many related parameters.

ABC variable selection

Posted in Books, Mountains, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , on July 18, 2018 by xi'an

Prior to the ISBA 2018 meeting, Yi Liu, Veronika Ročková, and Yuexi Wang arXived a paper on relying ABC for finding relevant variables, which is a very original approach in that ABC is not as much the object as it is a tool. And which Veronika considered during her Susie Bayarri lecture at ISBA 2018. In other words, it is not about selecting summary variables for running ABC but quite the opposite, selecting variables in a non-linear model through an ABC step. I was going to separate the two selections into algorithmic and statistical selections, but it is more like projections in the observation and covariate spaces. With ABC still providing an appealing approach to approximate the marginal likelihood. Now, one may wonder at the relevance of ABC for variable selection, aka model choice, given our warning call of a few years ago. But the current paper does not require low-dimension summary statistics, hence avoids the difficulty with the “other” Bayes factor.

In the paper, the authors consider a spike-and… forest prior!, where the Bayesian CART selection of active covariates proceeds through a regression tree, selected covariates appearing in the tree and others not appearing. With a sparsity prior on the tree partitions and this new ABC approach to select the subset of active covariates. A specific feature is in splitting the data, one part to learn about the regression function, simulating from this function and comparing with the remainder of the data. The paper further establishes that ABC Bayesian Forests are consistent for variable selection.

“…we observe a curious empirical connection between π(θ|x,ε), obtained with ABC Bayesian Forests  and rescaled variable importances obtained with Random Forests.”

The difference with our ABC-RF model choice paper is that we select summary statistics [for classification] rather than covariates. For instance, in the current paper, simulation of pseudo-data will depend on the selected subset of covariates, meaning simulating a model index, and then generating the pseudo-data, acceptance being a function of the L² distance between data and pseudo-data. And then relying on all ABC simulations to find which variables are in more often than not to derive the median probability model of Barbieri and Berger (2004). Which does not work very well if implemented naïvely. Because of the immense size of the model space, it is quite hard to find pseudo-data close to actual data, resulting in either very high tolerance or very low acceptance. The authors get over this difficulty by a neat device that reminds me of fractional or intrinsic (pseudo-)Bayes factors in that the dataset is split into two parts, one that learns about the posterior given the model index and another one that simulates from this posterior to compare with the left-over data. Bringing simulations closer to the data. I do not remember seeing this trick before in ABC settings, but it is very neat, assuming the small data posterior can be simulated (which may be a fundamental reason for the trick to remain unused!). Note that the split varies at each iteration, which means there is no impact of ordering the observations.