Archive for Jouy-en-Josas

SEINE AI

Posted in pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , on March 23, 2026 by xi'an

Ten days ago I took part in the SEINE AI 2026 workshop in Jouy-en-Josas, near Paris (homestead of HEC), organised by the Huawei Paris Research Center.. In which I was invited to speak, even though I felt sort of an outlier given the deeply machine-learning, entreprenarial orientation of the meeting, with its theme being Building the Agentic Future of ICT, given that I chose to present our most recent Bayesian adversarial privacy paper. Hence, I stood within a game-theoretic, Bayesian, formal landscape, presumably loosing most of the audience and keeping them away from their lunch!

Other speakers included Simon Lucas from Queen Mary London on Simulation-based AI, which I had trouble distinguishing from building a statistical model by goodness of fit (and using bandits used for update), while focussing on competing on some computer game challenges. And Volker Tresp from LMU München on a tensor brain model that he opposes to a Bayesian brain (with a related paper entitled Bayes or Heisenberg: Who(se) rules? which we discussed in general terms over lunch, namely Bayesian learning vs. quantum updating. And Michal Valko from INRIA Paris (and other companies), who went full blast against the Bradley-Terry model!, with a title of Nash and Nemirovski walk into a bar! With a half-time technique approximating Nash equilibria that reminded me of leapfrog. Much entertaining talk that further provided a game-theoretic transition to mine’s.

As an aside, I played yesterday with ChatGPT composing my talk slides out of our arXiv document and it proved a disaster, with hallucinations of results and concepts not in the paper and a complete mess of handling graphs, first creating generic, fake, unrelated pictures, then inserting actual graphs haphazardly throughout the slides. The sorry result I obviously did not use as the workshop did not seem the ideal place for this sort of prank! The actual version only recycles a few of its summarising slides. (With ye Norse farce proper colour choice!)

 

differentially private distributed Bayesian linear regression with MCMC

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , on August 30, 2023 by xi'an

An ICML 2023 paper by Barıs¸ Alparslan, Sinan Yıldırım¸ and Ilker Birbil that (re)addresses the issue of privacy when running a Bayesian regression analysis. Resorting to the common notion of differential privacy, imposing a limited variability if a single observation is modified, and a Gaussian randomisation of the observations.

“A differentially private algorithm constrains the difference between the probability distributions of the output values obtained from neighbouring data sets”

In the super classical setup of simple Normal linear regression, y=Xθ+σε. Summary statistics are chosen as

S=X’X and z=X’y,

(why the separation?) then randomised. (Keeping Ŝ definite positive? Not necessarily, it appear.) Inspired directly from Dwork & al. (2014). The authors still manage to spend an entire column in (re)deriving the conditional Normal distribution of z conditional on S and (θ,σ)… Which is later exploited for integrating z out in the MCMC algorithm.

“some important differences between our work and that of Bernstein & Sheldon (2019) [stem] from the choice of summary statistics and the consequent hierarchical structure used for modelling linear regression [and]lead to significant differences in the inference methods as well as significant computational advantages [O(d³) vs. O(d⁶)]”

In a distributed setting several agents are handling their own data and keep their privacy by the same mechanishttps://www.slideshare.net/xianblog/discussion-of-icml23pdfm [as in the top graph from the paper]. On principle, a Bayesian analysis of the resulting hierarchical model should directly consider the posterior on the global parameter by considering the distributions of the randomised pairs (ẑ,Ŝ). The elephant in the room is the distribution of the regressors, which is customarily unknown and not accounted for in a traditional Bayesian analysis. It is needed here due to the division in S and z, plus the randomisation step that calls for the posterior distribution of S given Ŝ. Elephant that is exfiltrated by either assuming Normality or substituting Ŝ for S without accounting for the noise! Definitely not exactly Bayesian. Another column is spent on the Metropolis-within-Gibbs simulation of the posterior…

Overall, I remain reserved about this approach, since it does not follow a clear Bayesian pathway and in particular does not incorporate privacy as part of the Bayesian decision analysis.

ellis unconference [not in Hawai’i]

Posted in pictures, Running, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on July 26, 2023 by xi'an

As ICML 2023 is happening this week, in Hawai’i, many did not have the opportunity to get there, for whatever reason, and hence the ellis (European Lab for Learning {and} Intelligent Systems] board launched [fairly late!] with the help of Hi! Paris an unconference (i.e., a mirror) that is taking place in HEC, Jouy-en-Josas, SW of Paris, for AI researchers presenting works (theirs or others’) presented at ICML 2023. Or not. There was no direct broadcasting of talks as we had (had) in CIRM for ISBA 2020 2021. But some presentations based on preregistered talks. Over 50 people showed up in Jouy.

As it happened, I had quite an exciting bike ride to the HEC campus from home, under a steady rain, crossing a (modest) forest (de Verrières) I had never visited before, despite it being a few km from home, getting a wee bit lost, stopped by a train Xing between Bièvre and Jouy, and ending up at the campus just in time for the first talk (as I had not accounted for the huge altitude differential). Among curiosities met on the way, “giant” sequoias, a Tonkin pond, Chateaubriand’s house.

As always I am rather impressed by the efficiency of AI-ML conferences run, with papers+slides+reviews online, plus extra material as in this example. Lots of papers on diffusion models this year, apparently. (In conjunction with the trend observed at the Flatiron workshop last Fall.) Below are incoherent tidbits from the presentations I attended:

  • exponential convergence of the Sinkhorn algorithm by Alain Durmus and co-authors, with the surprise occurrence of a left Haar measure
  • a paper (by Jerome Baum, Heishiro Kanagawa, and my friend Arthur Gretton) on Stein discrepancy, with an Zanella Stein operator relating to Metropolis-Hastings/Barker since it has expectation zero under stationarity, interesting approach to variable length random variables, not a RJMCMC, but nearby.
  • the occurance of a criticism of the EU GDPR that did not feel appropriate for synthetic data used in privacy protection.
  • the alternative Sliced Wasserstein distance, making me wonder if we could optimally go from measure μ to measure ζ using random directions or how much was lost this way.

\mathbb E[y|X=x] = \mathbb E\left[y\frac{f_{XY}(x,y)}{f_X(x)f_Y(y)}|X=x\right] = \frac{\mathbb E\left[y\frac{f_{XY}(x,y)}{f_Y(y)}|X=x\right]}{f_X(x)}

as (a) densities are replaced with kernel estimates, (b) the outer density may be very small, (c) no variance assessment is provided.

  • Markov score climbing and transport score climbing using a normalising flow, for variational approximation, presented by Christian Naesseth, with a warping transform that sounded like inverting the flow (?)
  • Yazid Janati not presenting their ICML paper State and parameter learning with PARIS particle Gibbs written with Gabriel Cardoso, Sylvain Le Corff, Eric Moulines and Jimmy Olsson, but another work with a diffusion based model to be learned by SMC and a clever call to Tweedie’s formula. (Maurice Kenneth Tweedie, not Richard Tweedie!) Which I just realised I have used many times when working on Bayesian shrinkage estimators

permanent position for research on computational statistics and “omics” data

Posted in pictures, Statistics, Travel, University life with tags , , , , , , , , on February 4, 2019 by xi'an

There is an opening at the French agronomy and genetics research centre, INRA, for a permanent research position on the country campus of Joyu-en-Josas, south-west of Paris, with focus on computational statistics (incl. machine-learning) and collaborations on omics data. The deadline is March 4. (The procedure is somewhat involved, as detailed in the guide for candidates.) I want to stress this is a highly attractive position in terms of academic surroundings (research only campus, nearby Paris=Saclay and Orsay campuses), of location (Paris in the fields), and of status since permanent really means permanent!