I only recently read the (foundational!) paper on Foundations of Bayesian learnin from synthetic data by Harrison Wilde, Jack Jewson, Sebastian Volmer [all associated with Warwick at some point] and [my long time friend] Chris Holmes, that merges Bayesian inference with differential privacy constraints thru generalised / Gibbs interface. Recouping with the M-open perspective in order to accommodate the misspecified nature of synthetic data. I like the approach very much in that it intersects a lot with my own views, excepts for following the differential privacy formalism. I however think that further progress could be made by adopting an even more Bayesian position.
Their key messages from that paper are that
- learning from synthetic data may prove damaging to your (data) health
- robustness unsurprisingly reduces the odds or magnitude of the damage
- real data can still be used to some extent
Since the (synthetic) generating model can be a GAN, the privacy requirement is such that noise is “injected” in the input data and in the learning mechanism. This is not discussed in the paper but highly conservative constraints surely make the DGP loose several learning points.
On the side, I also like the alternative of opposing data keeper and learner rather than data owner and adversary. Learning here means taking an optimal B decision about the actual data averaged over the true DGP. With a prior on the distribution of the actual data, while being unable to avoid misspecification in representing the (marginal) synthetic generation model.
Unsurprisingly, the alternative approach is relying on proper scoring rules as Bissiri et al. (2016). Rather than finding the distribution KL closest to the synthetic generative model, robustified by generalised B inference, either via downweighting or via ß-divergence. With a preference for the latter. Interestingly, the authors consider the optimal learning size for the synthetic data. Since bringing in more synthetic data does not mean better performances.
The main BayesComp meeting started right after the ABC workshop and went on at a grueling pace, and offered a constant conundrum as to which of the four sessions to attend, the more when trying to enjoy some outdoor activity during the lunch breaks. My overall feeling is that it went on too fast, too quickly! Here are some quick and haphazard notes from some of the talks I attended, as for instance the practical parallelisation of an SMC algorithm by Adrien Corenflos, the advances made by Giacommo Zanella on using Bayesian asymptotics to assess robustness of Gibbs samplers to the dimension of the data (although with no assessment of the ensuing time requirements), a nice session on simulated annealing, from black holes to Alps (if the wrong mountain chain for Levi), and the central role of contrastive learning à la Geyer (1994) in the GAN talks of Veronika Rockova and Éric Moulines. Victor Elvira delivered an enthusiastic talk on our massively recycled importance on-going project that we need to complete asap!
In a not-solely-ABC session, I appreciated Sirio Legramanti speaking on comparing different distance measures via Rademacher complexity, highlighting that some distances are not robust, incl. for instance some (all?) Wasserstein distances that are not defined for heavy tailed distributions like the Cauchy distribution. And using the mean as a summary statistic in such heavy tail settings comes as an issue, since the distance between simulated and observed means does not decrease in variance with the sample size, with the practical difficulty that the problem is hard to detect on real (misspecified) data since the true distribution behing (if any) is unknown. Would that imply that only intrinsic distances like maximum mean discrepancy or Kolmogorov-Smirnov are the only reasonable choices in misspecified settings?! While, in the ABC session, Jeremiah went back to this role of distances for generalised Bayesian inference, replacing likelihood by scoring rule, and requirement for Monte Carlo approximation (but is approximating an approximation that a terrible thing?!). I also discussed briefly with Alejandra Avalos on her use of pseudo-likelihoods in Ising models, which, while not the original model, is nonetheless a model and therefore to taken as such rather than as approximation.
After the (X country skiing) break, Lorenzo Pacchiardi presented his 
