Archive for federated learning

Information Geometry, Privacy and Monte Carlo workshop, ISM, 4-5 July 2026

Posted in Mountains, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , on July 6, 2026 by xi'an

After the (exciting) variety and spread of ISBA²⁶, here we are at much more focussed (and single-track), if equally exciting, workshop at the ISM. (With many participants from ISBA²⁶.)

On Saturday afternoon, Ajay Jasra talked about Particle filtering for state-space models with low, degenerate noise, with specific measure issues I did not really get, since the manifold attached to the noise was known, but the projected density may prove a challenge. Manifolds were also central to Kenji Fukumizu’s talk on Learning manifold structure and density with score-based models learning scores as projectors to the manifold, although it was unclear to me how this was possible when the manifold is unknown. Christophe Andrieu presented Geometry informed selection in multiple proposal MCMC which stems from an early multi-proposal (1998) proposal by Radford Neal and uses a multivariate ranking procedure to quantify a measure of surprise for the current  Markov chain value within the proposed ones. The crux for the efficiency of the approach may be in the choice of this ranking procedure. And Maria De Iorio talked about Efficient MCMC via similarity-driven proposals for discrete support targets, with similarities with ABC.

Completed with a human sized poster session where I reconnected with Spanish friends I had not seen for ages (by missing OBayes meetings).

On (pleasantly rainy) Sunday morning, Federica Milinanni detailled her Rapid mixing of stereographic MCMC for heavy-tailed sampling, essentially the same content as in Nagoya last FRiday, with a novel sub-Cauchy projection supposed to explore heavy tails better: while the regular stereographic projection turns the t-distribution with d degrees of freedom into a Uniform on the hypersphere, a sub-Cauchy projection turns the Cauchy into this uniform. In the privacy session I organised, Hongsheng Dai, member of our ERC Synergy project, presented an Online federated learning framework for classification, using DP as a criterion and achieving by adding noise to the loss function at each occurrence of the data production. Surprisingly increasing with the number of occurrences, not so much since the objective function keeps calling

Stefano Favaro described his Bayesian nonparametric privacy-preserving synthetic data generation method (for discrete data) that connects privacy protection and information preservation. (Incidentally I was unaware of the σ parameter of the Pitman-Yor process, which allows for a finite support when σ<0, but I cannot fathom the appeal of this extension, given the complete lack of connection between the positive and negative cases.) Surprisingly, non-parametric prediction does worse in terms of privacy, if not so surprising with discrete data since the predictive actually put weight on every datapoint. Resorting to  mechanism informativity by Wasserman and Zhou (2010) (with a surprise mention of my friend Arnaud Guilin!). And Joshua Bon gave his Persuasive Privacy talk of last Tuesday  (to be re-repeated two days later at ICML²⁶ in Seoul!). Except for changing the audience game from croissants to sumo wrestlers! (What will it be in Seoul!?) And adding much more details on the foundational elements of persuasive privacy.

The poster session was similarly enjoyable, even though I did not manage to get through all posters.

mostly Monte Carlo [20/02]

Posted in Statistics with tags , , , , , , , , , , , , , , , , on February 16, 2026 by xi'an

A new episode of our mostly Monte Carlo seminar, very soon coming near you (if in Paris):

On Friday 20/02/26, from 3-5pm at PariSanté Campus

15h: Paul Mangold (École Polytechnique, Palaiseau)

Convergence and Linear Speed-Up in Stochastic Federated Learning

In federated learning, multiple users collaboratively train a machine learning model without sharing local data. To reduce communication, users perform multiple local stochastic gradient steps that are then aggregated by a central server. However, due to data heterogeneity, local training introduces bias. In this talk, I will present a novel interpretation of the Federated Averaging algorithm, establishing its convergence to a stationary distribution. By analyzing this distribution, we show that the bias consists of two components: one due to heterogeneity and another due to gradient stochasticity. I will then extend this analysis to the Scaffold algorithm, demonstrating that it effectively mitigates heterogeneity bias but not stochasticity bias. Finally, we show that both algorithms achieve linear speed-up in the number of agents, a key property in federated stochastic optimization.

16h: Alain Durmus (École Polytechnique, Palaiseau)

A Mixture-based Framework for Guiding Diffusion Models

Inverse problems—such as image restoration from noisy or incomplete measurements and musical source separation—are ill-posed, making Bayesian approaches with learned generative priors especially appealing. Diffusion models provide powerful priors, but existing posterior sampling methods often rely on crude likelihood-gradient approximations and heavy task-specific tuning. In this talk, I will introduce a novel principled approach specifically designed to overcome these limitations. The core contribution of this approach is the construction of a mixture approximation of intermediate posterior distributions defined by the diffusion model. The sampling is carried out sequentially via Gibbs sampling, a Markov Chain Monte Carlo method, using a careful data augmentation scheme. Gibbs sampling is employed here due to its simplicity and theoretical guarantees, allowing for exact conditional updates at each iteration, thus ensuring stability and efficiency. One key advantage of the presented algorithm is its flexibility: it adapts to varying levels of computational resources by adjusting the number of Gibbs iterations. Consequently, substantial performance gains can be achieved by increasing inference-time computational effort. I will present extensive experimental results demonstrating empirical performance across diverse image restoration tasks, involving both pixel-space and latent-space diffusion models, and showcase its successful application in musical source separation.

 

Data science ethics [book review]

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on May 5, 2025 by xi'an

Data science ethics (concepts, techniques and cautionary tales), by David Martens, was published in 2022 by Oxford University Press. The book is inspired by the author’s  course on Data Science and ethics he has been teaching at the University of Antwerp. (With a link to his slides.) The 255p book proceeds by decomposing the ethics of data science into its different steps: data gathering (Chap. 2), data preprocessing (Chap. 3), modelling (Chap. 4), evaluation (Chap. 5), and deployment (Chap. 6). Following the `FAT Flow Framework´, where FAT stands for fairness, accountability, and transparency.

Do not expect much maths, stats, or anything quantitative: this book is mostly about concepts, even though some (mostly well-known) illustrations are provided. Chapter 2 includes a description of encryption (with homomorphic encryption treated in Chapter 4). And somewhat improbably quantum computing. Differential privacy gets a few pages with not a single formula (until Chapter 4, again).

Chapter 3 covers k-anonymity, record linkage, reidentification, (through cautionary tales) and discrimination through biases in the (learning) dataset.  Chapter 4 is defining ε differential privacy with the Laplace randomization as a possible implementation and with no critical stance on the limitations of the concept. The computation limitations of homomorphic encryption are more clearly pointed out. Federated learning is only quickly mentioned. The section about measuring fairness and reducing bias implies that some prior knowledge is available about whom is potentially discriminated and which covariates to add to the model. The last section on explicability of predictions is worthwhile in signalling the difficulty with most (black box) AI but the example opposing an SVM model to a logistic model is not tremendously convincing in that neither model is true.

Chapter 5 addresses the crucial challenge of ethical evaluation in a rather verbose and vague manner. For instance, with no instruction on how to resist adversarial attacks. Or criticising p-hacking and multiple testing while missing the elephant in the room (p-values!). Drifting from the topic when discussing the misdeeds of Diederik Stapel. Most of the same goes about Chapter 6 and its take on ethical deployment, when going through examples such as Google’s policies in China. Or general musing on the impact of AI on societal inequalities (with mentions of companies and CEOs who have since then back-pedalled on their ethical engagement). These chapters are lacking in tools and (more) practical recommendations.

One interesting aspect of the book is the attention paid to the EU(ropean) aspect of these concerns, through the GDPR (General DAta Protection Regulations) adopted by the European Parliament in 2016. (There is also a brief mention of China’s regulations, but no details beyond a reference. Maybe the Chinese edition differs.)

JSM 2024, Portland, miniday 4

Posted in Books, pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , on August 11, 2024 by xi'an

Final (half)day at JSM is always a sad thing as most people have left, people are busy dismantling booths and packing boxes, and the few remaining participants are fidgety and sitting on their suitcases (there may have been more suitcases than people in the main area that day!), cafés are minimally staffed or simply closed. Hence not the best time-slot to deliver one’s talk! Still, a few dozen people attended our session. The session topic was Bridge the Gap: Differential Privacy and Statistical Analysis, organised by Bei Jiang whom I met last summer in Kelowna, at a BIRS workshop on privacy. Where I spoke on setting up a complete decision-theoretic framework, as developed (and still in development) within our Ocean group, esp. Joshua Bon, Stan du Ché, and Judith Rousseau. (Rather than on the original plan of talking about convergence versus privacy, as a criticism of differential privacy.) The other talks were by Shurong Li, strongly set with differential privacy when record linkage is present, and Xuan Bi on a fully decentralised federated learning with local exchanges of global gradients that limit privacy leaks. (With a mention of gossip learning I hadn’t seen previously!)

I enjoyed even more the session due to Naisyin Wang giving a discussion on the three talks, in  closes the particular because I had not seen her in years, if a few times since she was my teaching assistant in Cornell in a Bayesian decision theory class I was building on the spot (with Linda Zhao as a student!). Which nicely closes the loop given the topic of my talk, of which she was quite supportive! While calling for debiasing post-processing and wondering about a two-dimensional decision theoretic perspective rather than a unidimensional one under hard privacy constraints.

On the way out, after parting from the few friends remaining in the convention centre, I spotted the nearby (steel) bridge being raised, although I could not see the boat responsible for it. I had been unaware of this possibility while running over and under it, as well as swimming thrice under it. And we left Portland in the early afternoon, heading for Seattle and a celebration of Adrian Raftery’s career.

ritorno sereno in Serenissima

Posted in pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , , on June 29, 2024 by xi'an