Archive for minimaxity

optimal sampling for kernel quadrature on unbounded domains

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , on May 22, 2026 by xi'an

My PhD student Edoardo Bandoni, along with Julien Stoehr and myself, completed a paper on validating (Bayesian) kernel quadrature with unbounded domains of integration. Which connects with probabilistic numerics, since the integrand is modelled as a Gaussian process. And RKHS methods. As opposed to Monte Carlo estimators, quadrature methods approximate integrals of smooth functions with worst-case error decaying at a minimax rate α/d for smoothness α in dimension d. Existing rate-optimal quadrature methods often depend on deterministic point sets tailored to a specific kernel, making them sensitive to misspecification and thus less robust in practice. This paper studies instead randomised quadrature methods, with a focus on robustness rather than on kernel-specific optimality. We construct an explicit, n-dependent, sampling distribution that achieves minimax rates for worst-case errors over smoothness classes without requiring knowledge of the kernel. This kernel-agnostic design does improve robustness while retaining optimal rates and extends Briol et al.  (2019) to the unbounded case. Which cannot always be easily handled by a change of variables. Our result thus mostly covers unbounded sampling measures such as Gaussian and Student-t distributions, extending beyond compact domains. The results provide both theoretical guarantees and a practical recipe for robust, rate-optimal, randomised quadrature. [The above is mostly stated in the abstract.]

a (sunny, crisp) day at ICSDS 2025

Posted in pictures, Running, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , on December 19, 2025 by xi'an

While my first day at ICSDS 2025 was somewhat hectic, having realised late the night before that I was giving a talk!—I had forgotten I had submitted a title at registration time and never received any communication from the organisers, including (or excluding) a request for an abstract. I thus hastily updated my November talk in Sevilla for my December talk in Sevilla! but paid less attention than needed to the sessions I attended—, Wednesday was more peaceful—esp. after a 16K run along the Guadalquivir—and I engaged into two great Bayesian learning sessions, one that seemed designed for me!, involving my (40y long friend) Ed George on his latest result on proper prior minimaxity and shrinkage, with our late friend Bill Strawderman as a co-author since they worked on the problem prior to Bill’s demise, Charles Margossian on variational inference preserving some symmetries in the target and hence keeping the same statistics, with elliptically symmetric families, and Fletcher Christensen on DIC for some mixed models, with references to our “DIC’s eights” paper (but still picking one version of DIC in the end!)

The second session was on prediction learning!—with me as the chair, as I realized one minute before! AI !—with (my friend) Veronika Rockova using AI predictions as a prior predictive and connecting them with Bayesian nonparametrics, Kenyon Ng (who visited me last Spring) on a similar approach using pretrained transformers like TabPFN and martingale posterior inference, Lorenzo Cappello in a generalisation of martingale prediction and Andrea Ghiglietti on the mathematics of an involved urn system.


The afternoon session was a plenary talk by Daniela Witten in the magnificent building of the Real Fabrica de Tabacos, but the room was unfortunately too small for the audience and I could not enter. Hopefully her talk will have a significant intersection with the CRiSM colloquium she delivers in Warwick late January. I thus walked around the old town till the following poster session, held in the Real Fabrica courtyard, under the sun. As I got involved into a deep discussion of the relevance of mirror meetings (which I defend!) versus the dangers on principal (parent) conferences (which can be mitigated by the mirror conference participants registering, to some extent, for the principle one)—more to come on the ‘Og and in the ISBA Bulletin!—, I did not peruse the available posters, sorry…

And, by the way, the conference organisers also revealed the location of ICSDS 2026 which is Croatia, my first bet! In the city of Split we visited in 2023.

William (Bill) Strawderman (1941-2024)

Posted in pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , on October 3, 2024 by xi'an

Earlier today, I was informed by several of our mutual friends that my long-time friend Bill Strawderman had sadly passed away yesterday, after fighting a cancer for the past months. I remember quite clearly meeting Bill in the Fall of 1988 in front of White Hall, which hosted the Cornell maths department at the time, as he was visiting George Casella from Rutgers where he spent most of his career. I was most eager to meet him as I had worked on several of his landmark papers during my PhD on shrinkage estimation, as well as a bit impressed. But his kindness, modesty, and congenial personality quickly put me at ease and we spent the rest of his visit discussing shrinkage but also literature and music. Especially Dickens! After that we met and collaborated quite regularly, to the point he started visiting France upon my return, at Paris 6 (Pierre & Marie Curie) University first, and then in Rouen, where he became a adjunct professor and launched a life-long collaboration and friendship with Dominique Fourdrinier. As my interest in shrinkage estimation dwindled along the years, we did not keep collaborating for the past two decades, but we remained in touch and I was very happy to participate in his 80th anniversary celebration in Rutgers two years ago. His contributions to the field are notable and several papers of his were part of the Bayesian classics I was giving my graduate class a few years ago. From the fabulous minimaxity paper of 1984, along with George Casella, to admissible estimators dominating the positive-part James-Stein estimator, to sufficient conditions of minimaxity for proper Bayes estimators, to decision theoretic properties of Bayesian credible interval estimators, to loss estimation, not to mention his more applied side… Besides his fabulous sense of humour, which made many evenings with him memorable, I will also cherish the memory of a bon vivant who liked good food and good wines, incl. the Calvados apple brandy I would bring him at each of my visits.

safe Bayes & e-values & least favourable priors

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , , , , , , , on March 2, 2024 by xi'an

The paper by Peter Grünwald, Rianne de Heide and Wouter Koolen on safe testing was read before The Royal Statistical Society at a meeting organized by the Research Section on Wednesday, 24th January, 2024, after many years in the making, to the point that several papers based on this initial one have appeared in the meanwhile, incl. some submissions to Biometrika. Like this one in the current issue of Statistical Science dedicated to reproducibility and replicability. Joshua Bon and I wrote a discussion that synthesised the following and sometimes rambling remarks.

Overall, this is a mind-challenging paper with definitely original style and contents for which the authors are to be congratulated!

“…p-values are interpreted as indicating amounts of evidence against the null, and their definition does not need to refer to any specific alternative H¹. Exactly the same holds for e-values: the basic interpretation ‘a large e-value provides evidence against H⁰’ holds no matter how the e-variable is defined, as long as it satisfies (1). If they are defined relative to H¹ that is close to the actual process generating the data they will grow fast and provide a lot of evidence, but the basic interpretation holds regardless.”

About the entry section, one may ask why would a Bayesian want to test the veracity of a null hypothesis. The debate has been raging since the early days, although Jeffreys spent two chapters of his book on the topic of testing. (While appearing in Example 5 p.14 for his point estimation prior.) From an opposite viewpoint, the construction of e-values and such in the paper is highly model dependent, but all models are wrong! and more to the point both hypotheses may turn out to be wrong for misspecified cases. The notion thus seems on the opposite to be very M-close, with no idea of what is happening under misspecified models or why is rejecting H⁰ the ultimate argument.

When introducing e-values, (1) is not a definition per se, since otherwise E≡1 would be an e-value. This is unfortunate as the topic is already confusing enough. E[E] must be larger than 1 under H¹, otherwise product of e-values would always degenerate to zero (?)

The points

  1. behaviour under optional continuation [by a martingale reasoning]
  2. interpretation as ‘evidence against the null’ as gambling [unethical!]
  3. in all cases preserving frequentist Type I error guarantees
  4. e-variables turn out to be Bayes factors based on the right Haar prior [rather than sometimes with highly unusual (e.g. degenerate) priors? p.4]
  5. e-variables need more extreme data than p-values in order to reject the null

are rather worthwhile, even though 2. is vague and 3. is firmly frequentist. Any theory involving Haar priors (and even better amenability) cannot be all wrong, though, even considering that Haar priors are improper. The optional continuation in 1. is a nice argument from a Bayesian viewpoint since it has also been used to defend the Bayesian approach. Point 4. brings a formal way to define least favourable priors in the testing sense. One may then wonder at the connection with the solution of Bayarri and Garcia-Donato (Biometrika, 2007). The perspective adopted therein is somehow an inverse of the more common stance when the prior on H⁰ is the starting point [and obviously known]. So, is there any dual version of e-values where this would happen, i.e. leading to deriving the optimal prior on H¹ for a given prior on H⁰? (Which would further offer a maximin interpretation.) Theorem 1 indeed sounds like the minimax=maximin result for test settings. (In Corollary 2, why is (10) necessarily a Bayes factor, given the two models?)

While I first thought that the approach leads to finding a proper prior, the “Almost Bayesian Case” [p.17] (ABC!!) comes to justify the use of a “common” improper prior over nuisance parameters under both hypotheses, which while more justifiable than in the original objective Bayes literature, remains unsatisfactory to me. But I like the notion in 2.2 [p.10] that a prior chosen on H¹ forces one to adopt a particular corresponding prior on H⁰, as it defines a form of automated projection that we also considered in Goutis [RIP] and Robert (Biometrika, 1998). Corollary 2 is most interesting as well. However, taking the toy example of H⁰ being a normal mean standing in (-a,a) seems to lead to the optimal prior on H⁰ being a point mass at +/- a for any marginal m(y) centred at zero. Which is a disappointing outcome when compared with the point mass situation. It is another disappointment that the Bayes Factor cannot be an e-value since (6) fails to hold, but (1) is not (6) and one could argue that the BF is an e-value when integrating under the marginals!

As a marginalia, the paper made me learn about the term (and theme) tragedy of the commons, a concept developed by [the neomalthusian and eugenist] Garett Hardin.

In conclusion, we congratulate the authors on this endeavour but it remains unclear to us (as Bayesians) (i) how to construct the least favourable prior on H0 on a general basis, especially from a computational viewpoint, and, more importantly, (ii) whether it is at all of inferential interest [i.e., whether it degenerates into a point mass]. With respect to the sequential directions of the paper, we also wonder at the potential connections with sequential Monte Carlo, for instance, towards conducting sequential model choice by constructing efficiently an amalgamated evidence value when the product of Bayes factors is not a Bayes factor (see Buchholz et al., 2023).

probabilistic numerics [book review]

Posted in Books, pictures, Statistics, Travel with tags , , , , , , , , , , , , , , , , , , , , on July 28, 2023 by xi'an

Probabilistic numerics: Computation as machine learning is a 2022 book by Philipp Henning, Michael Osborne, and Hans Kersting that was sent to me by CUP (upon my request and almost free of charge, as I had to pay custom charges, thanks to Brexit!). With the important message of bringing statistical tools to numerics. I remember Persi Diaconis calling for (such) actions in the 1980’s (and even reading a paper of his on the topic along with George Casella in Ithaca while waiting for his car to get serviced!).

From a purely aesthetic view point, the book reads well, offers a beautiful cover and sells for a quite reasonable price for an academic book. Plus it is associated with a website containing draft version of the book. Code, links to courses, research, conferences are also available there. Just a side remark that it enjoys very wide margins that may have encouraged an inflation of footnotes (but also exercises). Except when formulas get in the way (as e.g. on p.40).

The figure below is an excerpt from the introduction that sets the scene of probabilistic numerics involving algorithms as agents, gathering data and making decisions, with an obvious analogy with standard Bayesian decision theory. Modelling uncertainty missing from the picture (if not from the book, as explained later by the authors as an argument against attaching the label Bayesian to the field). Also referring to Henri Poincaré for the origination of the prior vs posterior uncertainty about a mathematical quantity. Followed by early works from the Russian school of probability, somewhat ignored until the machine-learning revolution and a 2012 NIPS workshop organised by the authors. (I participated to a follow-up workshop at NIPS 2015.)

In this nicely written section, I have an objection to the authors’ argument that a frequentist, as opposed to a Bayesian, “has the loss function in mind from the outset” (p.9), since the loss function is logically inseparable from the prior and considered from the onset. I also like very much the conclusion to that introduction, namely that the main messages (from the book) are that (verbatim)

  • classical methods are probabilist (p.10)
  • numerical methods are autonomous agents (p.11)
  • numerics should not be random (if not a rejection of the concept of Monte Carlo methods, p.1, but probabilistic numerics being opposed to stochastic numerics, p.67)
  • numerics must report calibrated uncertainty (p.12)
  • imprecise computation is to be embraced (p.12)
  • probabilistic numerics consolidates numerical computation and statistical inference (p.13)
  • probabilistic numerical algorithms are already adding value (p.13)
  • pipelines of computation demand harmonisation

“Is it still reasonable to labour under computational constraints conceived in the 1940s?” (p.113)

“rather than being equally good for any number of dimensions, Monte Carlo is perhaps better thought of as being equally bad” (p.110)

Chapter I is a 40p infodump (!) on mathematical concepts needed for the following parts. Chapter II is about integration, opposing again PN and Monte Carlo (with strange remark that MCMC does not achieve √N convergence rate, p.72). In the sense that the later is frequentist in that it does not use a prior [unless considering a limiting improper version as in Section 12.2, an intriguing concept in this setup as I wonder whether or not improper priors can at all be contemplated] on the object of interest and hence that the stochasticity does not reflect uncertainty but rather the impact of the simulated sample. Advocating Bayesian quadrature (with some weird convergence graphs exhibiting a high variability with the number of iterations that apparently is not discussed) and bringing in the fascinating perspective of model choice in that framework (leading to compute a posterior probability for each model!). Being evidently biased towards Monte Carlo, I find the opposition in Chapter 12 unnecessarily antagonistic, while presenting Monte Carlo methods as a form of minimax solution, the more because quasi-Monte Carlo methods are hardly discussed (or dismissed). As illustrated by the following picture (p.115) and the above quotes. (And I won’t even go into the absurdity of §12.3 trashing pseudo-random generators as “painfully dumb”.)

Chapter III is a sort of dual of Chapter II for linear algebra numerics, primarily solving linear equations by Gaussian solvers, which introduces new concepts like Krylov sequences, although it sounds quite specific (for an outsider like me). Chapters IV and V deal with the more ambitious prospect of optimisation. Reconsidering classics and expanding into Bayesian optimisation, using Gaussian process priors and defining specific loss functions. Bringing in a strong link with machine learning tools and goals. [citation typo on p.277]. Chapter VII addresses the resolution of ODEs by a Bayesian state space model representation and (again!) Gaussian processes. Reaching to mentioning inverse problems and offering a short finale on prospective steps for interested readers.

[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE.]