Archive for score function
OWABI⁷, 26 February 2026: Prequential posteriors (11am UK time)
Posted in Books, Statistics, University life with tags ABC, approximate Bayesian inference, diffusion model, generalised Bayesian inference, generative model, OWABI, prequential loss, reinforcement learning, score function, sequential ABC, sequential Monte Carlo, simulation-based inference, SMC, University of Warwick, webinar on February 25, 2026 by xi'anOWABI⁷, 29 January 2026: Sequential Neural Score Estimation (11am UK time)
Posted in Books, Statistics, University life with tags ABC, approximate Bayesian inference, diffusion model, generalised Bayesian inference, generative model, OWABI, score function, sequential Monte Carlo, simulation-based inference, University of Warwick, webinar on January 21, 2026 by xi'anveniSBA²
Posted in Books, pictures, Running, Statistics, Travel, University life with tags ABC, Bayesian conference, Bayesian data analysis, Bayesian decision theory, BDA, canal, coherence, composite likelihood, ISBA 2024, Italia, mixtures of experts, neural network, normalising flow, One World ABC Seminar, prior selection, score function, sequential Monte Carlo, simulation-based inference, SIS model, slaughterhouse, subjective prior, summary statistics, Università Ca' Foscari Venezia, Venezia, Venice on July 4, 2024 by xi'an
After another morning cycle of 2Xing Porte della Libertà (under a light and pleasant rain) and swimming in Sant’ Alviso (in too warm a water), I did not make it for the beginning of the Bayesian deep learning session, breakfast oblige!, and cumulated with different percolation events (ie, meeting friend after friend on my way to the classroom), I could not get enough of the session to report anything even barely useful!
As I did not rush fast enough to Andrew’s Foundation lecture (another sequence of percolations!), I had to stand in the back of the packed main amphitheatre (and former sorting hall of the Venice slaughterhouse!), Guido Cazzavillan’s Aula Magna, while he talked a fresco about some holes in Bayesian data analysis (the analysis, not the book!), those being [verbatim]
- the usual rules of conditional probability fail in the quantum realm,
- flat or weak priors lead to terrible inferences about things we care about,
- subjective priors are incoherent,
- Bayesian decision picks the wrong model,
- Bayes factors fail in the presence of flat or weak priors,
- for Cantorian reasons we need to check our models, but this destroys the coherence of Bayesian inference.
After lunch, I attended the (mostly sequential) simulation based inference (renamed from ABC!) session with a composite likelihood proposal by Lorenzo Rimella, that uses marginals to approximate the likelihood of a hidden Markov SIS epidemic model by composite likelihood towards getting more efficient if inexact versions. Then [1WABC webinar co-organiser] Umberto Picchini on surrogates for likelihood and posterior functions, with sequential improvements (w/o ABC and w/o neural networks). Called “Sequential mixture posterior and likelihood estimation”, using mixtures of experts when the weights are functions of the observed or simulated y. With adapting the number of components in the mixture. Comparing favourably with normalising flows. And Wentao Li on correcting by ABC for composite likelihood as in Ruli et al. (2016). Where a posterior distribution given composite scores (seen as [summary] statistics) is employed but requires a convergent estimator of the unknown parameter.
No congratulation today to our PhD student who managed to fall in a canal (but survived)..!

robust privacy
Posted in Books, Statistics, University life with tags #ERCSyG, Annals of Statistics, consistency, differential privacy, gradient descent, JASA, M-estimation, Newton-Raphson algorithm, Ocean, privacy, robustness, score function, stochastic gradient descent on May 14, 2024 by xi'an
During a recent working session, some Oceanerc (incl. me) went reading Privacy-Preserving Parametric Inference: A Case for Robust Statistics by Marco Avella-Medina (JASA, 2022), where robust criteria are advanced as efficient statistical tools in private settings. In this paper, robustness means using M-estimators T—as function of the empirical cdf—with basis score functions Ψ, defined as
where Ψ is bounded. A construction further requiring that one can assess the sensitivity (in Dwork et al, 2006, sense) of a queried function, sensitivity itself linked with a measure of differential privacy. Because standard robustness approaches à la Huber allow for a portion of the sample to issue from an outlying (arbitrary) distribution, as in ε-contaminations, it makes perfect sense that robustness emerges within the differential framework. However, this common sense perception does not seem good enough for achieving differential privacy and the paper introduces a further randomization with noise scaled by (n,ε,δ) in the following way
that also applies to test statistics. This scaling seems to constitute the central result of the paper, which establishes asymptotically validity in the sense of statistical consistency (with the sample size n). But I am left wondering whether this outcome counts as supporting differential privacy as a sensible notion…
“…our proofs for the convergence of noisy gradient descent and noisy Newton’s method rely on showing that with high probability, the noise introduced to the gradients and Hessians has a negligible effect on the convergence of the iterates (up to the order of the statistical error of the non-noisy versions of the algorithms).” Avella-Medina, Bradshaw, & Loh
As a sequel I then read a more recent publication of Avella-Medina, Differentially private inference via noisy optimization, written with Casey Bradshaw & Po-Ling Loh, which appeared in the Annals of Statistics (2023). Again considering privatised estimation and inference for M-estimators, obtained by using noisy optimization procedures (noisy gradient descent, noisy Newton’s method) and constructing noisy confidence regions, that output differentially private avatars of standard M-estimators. Here the noisification goes through a randomisation of the gradient step like
where B is an upper bound on the gradient Ψ, η is a discretization step, and K is the total number of iterations (thus fixed in advance). The above stochastic gradient sequence converges with high probability to the actual M-estimator in n and not in K, since the upper bound on the distance scales in √K/n. Where does the attached privacy guarantee come from? It proceeds by an argument of a composition of a sequence of differentially private outputs, all based on the same dataset.
“…the larger the number [K] of data (gradient) queries of the algorithm, the more prone it will be to privacy leakage.”
The Newton method version is a variation on the above stochastic gradient descent. Except it seems to converge faster, as illustrated above.

