
Archive for ERC
di ritorno a Venezia, nella privacy oceanica
Posted in pictures, Running, Statistics, Travel, University life with tags #ERCSyG, Bayesian privacy, Campo Sant'Alviso, canals, differential privacy, ERC, ERC Synergy Grant, Italia, laguna, Les Houches, Ocean, San Giobbe, swimming pool, uncertainty quantification, Università Ca' Foscari Venezia, Venezia, visiting position, workshop on March 15, 2026 by xi'an
mostly Monte Carlo [13/03]
Posted in Statistics, Travel, University life with tags #ERCSyG, Adam, ERC, Flatiron Institute, Gibbs sampler, Hamiltonian Monte Carlo, HMC, INRIA, Kantorovich semi-distances, Langevin diffusion, Markov kernel, Markov semigroup, MCMC, NUTS, Ocean, Paris, PariSanté campus, seminar, WALNUTS on March 10, 2026 by xi'an
A new episode of our mostly Monte Carlo seminar, very soon coming near you (if in Paris):
On Friday 13/02/26, from 3-5pm at PariSanté Campus
15h00: Pierre Del Moral (INRIA, Bordeaux)
On the Kantorovich contraction of Markov semigroup
We present a novel operator theoretic framework to study the contraction properties of Markov semigroups with respect to a general class of Kantorovich semi-distances, which notably includes Wasserstein distances. This rather simple contraction cost framework combines standard Lyapunov techniques with local contraction conditions. Our results can be applied to both discrete time and continuous time Markov semigroups, and we illustrate their wide applicability in the context of (i) Markov transitions on models with boundary states, including bounded domains with entrance boundaries, (ii) operator products of a Markov kernel and its adjoint, including two-block-type Gibbs samplers, (iii) iterated random functions and (iv) diffusion models, including overdampted Langevin diffusion with convex at infinity potentials.
16h00: Bob Carpenter (Flatiron Institute, New York)
GIST, WALNUTS, and Continuous Nutpie: mass-matrix and step-size adaptation for Hamiltonian Monte Carlo
I will introduce Gibbs self tuning (GIST), our new technique for coupling tuning parameters and conditionally Gibbs-sampling them per iteration in Hamiltonian Monte Carlo. Then I will turn to the within-orbit adaptive NUTS (WALNUTS) sampler, which adapts the step size every leapfrog step in order to conserve the Hamiltonian. Empirical evaluations on varying multi-scale target distributions, including Neal’s funnel and the Stock-Watson stochastic volatility time-series model, demonstrate that WALNUTS achieves substantial improvements in sampling efficiency and robustness. I will review the Nutpie mass-matrix adaptation scheme, which is designed to minimize Fisher divergence by estimating the mass matrix as the geometric midpoint (aka barycenter) between the inverse covariance of the draws and the covariance of the scores of the draws. Then I will describe a continuously adapting version that adapts per iteration by continuously discounting the past rather than updating in fixed blocks. I will also show how the Adam optimizer outperforms dual averaging for step-size adaptation. I will conclude by considering a lock-free multi-threading implementation that automatically monitors adaptation and sampling for convergence for automatic stopping.
Bayesian, adversarial, oceanic, privacy
Posted in Books, Statistics, University life with tags #ERCSyG, adversarial learning, Alice and Bob, Bayesian decision theory, Bayesian privacy, cryptography, data privacy, Data Protection Act, differential privacy, disclosure risk, ERC, European Research Council, ex ante risk., insufficient statistic, Montréal, Ocean, PariSanté campus, privacy-utility trade-off, release mechanism, RSA algorithm, statistical disclosure control, The Prairie Chair, Université Paris Dauphine, xkcd on March 6, 2026 by xi'anWe just arXived a new paper on Bayesian privacy! We meaning Cameron Bell, Antoine Luciano, Timothy Johnston and myself, as members of my ERC OCEAN lab at PariSanté and Paris Dauphine. While sharing the same ground as my recent paper with James Bailie, Joshua Bon and Judith Rousseau, this one is definitely more mainstream Bayesian in that the entire decision process falls under the Bayesian hat, with the ultimate decision being the choice of the release mechanism by the data holder (or hoarder!). To rationalise this decision process, we break the framework as resulting from the actions of three actors, namely the data holder, Alice, the data scientist, Bob, and the eavesdropper. Eve. (As in my earlier posts on solving Le Monde’s math puzzles, we could have used pronouns from other cultures, but I feared this would have confused some of the readers. Incidentally, I found out that the earliest use of the first two pronouns was within the groundbreaking cryptography 1977 paper of Rivest, Shamir and Adleman, bringing the RSA algorithm to the World! With Eve appearing in an early, highly-cited privacy paper by Montréal’s Bennett, Brassard, and (unconnected to me!) Robert, in 1988.)
We thus consider a Bayesian setting in which, given data x, held by Alice, inference is to be performed by Bob on a parameter θ. Performing such inference requires Alice releasing information derived from x, which may contain sensitive content, exploited by Eve. Our approach is to compare Alice’s release mechanisms according to both the quality of inference on θ (from Bob’s viewpoint) and the privacy leakage regarding x (sought by Eve and dreaded by Alice). To formalise this evaluation, we posit that Alice refers to a loss function that is a linear combination of Bob’s and Eve’s losses, the weight on Eve’s loss being then negative. (An alternative to be considered in future work is Alice using a ratio of Bob’s and Eve’s losses, possibly set to different powers, the rationale being that a zero loss for Eve is intolerable for Alice.) As in Bayesian experimental design, a prior on the data is necessary for Eve to infer on the hidden data based on the release mechanism and released output and for Alice to evaluate the risk of said release mechanism . (They may differ, as long as they are both made public.) To calibrate Alice’s loss, we opted for a balance that returns the same risk for a full data release and a total lack of release. In specific, informed, settings, other weights could be chosen. While finding the optimal release strategy is impossible but for highly discrete settings, the framework obviously allows for the ranking of natural strategies like insufficient statistics and synthetic datasets. Comments welcome!
mostly Monte Carlo [20/02]
Posted in Statistics with tags #ERCSyG, École Polytechnique, ERC, federated learning, generative prior, Gibbs sampling, inverse problems, machine learning model, MCMC, Ocean, Palaiseau, Paris, PariSanté campus, privacy, Scaffold algorithm, seminar, stochastic gradient on February 16, 2026 by xi'an
A new episode of our mostly Monte Carlo seminar, very soon coming near you (if in Paris):
On Friday 20/02/26, from 3-5pm at PariSanté Campus
15h: Paul Mangold (École Polytechnique, Palaiseau)
Convergence and Linear Speed-Up in Stochastic Federated Learning
In federated learning, multiple users collaboratively train a machine learning model without sharing local data. To reduce communication, users perform multiple local stochastic gradient steps that are then aggregated by a central server. However, due to data heterogeneity, local training introduces bias. In this talk, I will present a novel interpretation of the Federated Averaging algorithm, establishing its convergence to a stationary distribution. By analyzing this distribution, we show that the bias consists of two components: one due to heterogeneity and another due to gradient stochasticity. I will then extend this analysis to the Scaffold algorithm, demonstrating that it effectively mitigates heterogeneity bias but not stochasticity bias. Finally, we show that both algorithms achieve linear speed-up in the number of agents, a key property in federated stochastic optimization.
16h: Alain Durmus (École Polytechnique, Palaiseau)
A Mixture-based Framework for Guiding Diffusion Models
Inverse problems—such as image restoration from noisy or incomplete measurements and musical source separation—are ill-posed, making Bayesian approaches with learned generative priors especially appealing. Diffusion models provide powerful priors, but existing posterior sampling methods often rely on crude likelihood-gradient approximations and heavy task-specific tuning. In this talk, I will introduce a novel principled approach specifically designed to overcome these limitations. The core contribution of this approach is the construction of a mixture approximation of intermediate posterior distributions defined by the diffusion model. The sampling is carried out sequentially via Gibbs sampling, a Markov Chain Monte Carlo method, using a careful data augmentation scheme. Gibbs sampling is employed here due to its simplicity and theoretical guarantees, allowing for exact conditional updates at each iteration, thus ensuring stability and efficiency. One key advantage of the presented algorithm is its flexibility: it adapts to varying levels of computational resources by adjusting the number of Gibbs iterations. Consequently, substantial performance gains can be achieved by increasing inference-time computational effort. I will present extensive experimental results demonstrating empirical performance across diverse image restoration tasks, involving both pixel-space and latent-space diffusion models, and showcase its successful application in musical source separation.
persuasive (and Oceanic) privacy
Posted in Books, Mountains, pictures, Statistics, Travel, University life with tags #ERCSyG, adversarial strategy, arXiv, Bayesian decision theory, Bayesian privacy, decision-making agents, differential privacy, ERC, ERC Synergy Grant, fairness, French Alps, game theory, Les Houches, Ocean, uncertainty quantification, Università Ca' Foscari Venezia, Université Paris Dauphine, Venice, workshop on February 3, 2026 by xi'an
I am quite excited about the paper James Baillie, Joshua Bon, Judith Rousseau, and myself just arXived! A novel framework for measuring privacy we have been working on for at least the past year, partly through the previous Les Houches privacy workshops. In the spirit of these workshops and the larger scale ERC Synergy grant OCEAN, we develop therein a rather generic Bayesian game-theoretic perspective on achieving statistical privacy. It involves a Sender (observing the original data and delivering a limited output) and a Receiver (with potential adversarial intentions). The paper mostly focus on setting a theoretical framework, including the creation of new, purpose-driven privacy definitions that are rigorously justified, while also allowing for the assessment of existing privacy guarantees through game theory. While this was not our original intent, we show that pure and probabilistic differential privacy notions, in the Dwork et al. (2006) sense, are special cases of our framework. This setting provides new interpretations of the post-processing inequality. Furthermore, and somewhat more importantly, we also prove that our privacy guarantees can be established for deterministic algorithms, which are outside current privacy standards. Hopefully, we’ll make further progress at the incoming privacy workshop next month, to be held in Venice (again).
