Throughover the workshop in Chennai floated (!) the figure of Jean/André Ville, with his inequality generalising Markov’s, who invented martingales. He is not such a well-known figure in France—at least to me!—, despite having led a rather exceptional life, from being a visiting scholar in Berlin (in the Maison académique de Berlin, along with a certain Jean-Paul Sartre) and Vienna in the 1930s, to his wife being (in Berlin) one of the many (disposable and despised) lovers of JP Sartre (to whom an open-minded or clueless Ville later sent his thèse d’université on martingales and collectives, a much more substantial piece of work than the current PhD), to him working with German and Austrian mathematicians and logicians, such as Popper, Gödel, and Wald–who, what a coïncidence!, died in India from a plane crash in 1950 that had left from Chennai–and being impressed enough by the latter to passing an economics degree in the Sorbonne when back in Paris, establishing a minimax result for a zero-sum matrix game with two players, to his counter-example to von Mises’ kollectiv, to his nickname of the King of Counterexamples in the Viennese mathematics seminar, to him operating the first (Bull) computer at the Université de Paris. (Glenn Shafer wrote a detailed accounting of his youth, on which this post is based, up to his thesis defence but a few days from France mobilising for war–where his collegue Wolfgang Doeblin would kill himself the year after, to avoid capture–. With Bernard Bru, Edmond Malinvaud and Alain Trognon among the people who helped.) After the war, he worked several years as a prépa maths teacher before working for a French State electricity companion on signal theory and Monte Carlo methods, and then returning to Université de Paris as a professor in 1957.
Archive for martingales
André ou Jean Ville (1910-1989)
Posted in Books, pictures, Travel, University life with tags 25w5482, Abraham Wald, Alain Trognon, André Ville, Andrey Markov, École Normale Supérieure, Berlin, Bernard Bru, BIRS-CMI, Bull computers, Chennai, Edmond Malinvaud, Emile Borel, George Darmois, Glenn Shafer, history of Monte Carlo, history of statistics, India, Jean Ville, Jean-Paul Sartre, Karl Popper, Kurt Gödel, martingales, Maurice Fréchet, minimax strategy, Monte Carlo methods, Paris, Paul Lévy, plane crash, Richard von Mises, signal theory, Simone de Beauvoir, Sorbonne, supermartingale, two-player game, Université de Paris, Vienna, Ville's inequality, Wolfgang Doeblin, WW II on August 12, 2025 by xi'anon stopping rules
Posted in Books, Statistics with tags 25w5482, Bayesian testing, BIRS-CMI, Chennai, Chennai Mathematical Institute, e-values, evidence, game theory, inference, Likelihood Principle, Madras, martingales, p-values, sequential analysis, stopping rule, Tamil Nadu, The American Statistician, workshop on August 3, 2025 by xi'an
The workshop in Chennai and its focus on sequential procedures made me realise (among other things) I had never read Cornfield’s 1966 TAS paper on sequential testing and the likelihood principle:
“By sequential analysis I mean any form of analysis in which the conclusion depends not only on the data, but also on the stopping rule.”
Written with little maths and formalism, this paper argues that keeping a fixed critical level amounts to keeping a fixed amount of evidence. Hence constituting an early critique of p-values even though not expressed in such terms. The part of the paper related with the likelihood principle does not address testing or evidence in a Bayesian way. As a side (late awakening) remark, iid observations in sequential settings are not longer independent, conditional on the stopping rule realisation N=n, since they are constrained by the fact that the stopping rule realisation is n and not n-1, n-2, … For a short while, I thought it was in turn impacting the distribution of any “sufficient” statistic one may propose, with a normalising constant that depends on the unknown parameter and hence cannot be neglected. Over all those years, I had never though of the modification of sufficiency characteristics in such contexts. But in fine the pair made of the value of the stopping rule and of the unsequential sufficient statistics proves enough. And the normalisation constant is the probability that the stopping rule.. stops!, which is equal to one! For the same short while, I was then wondering that the stopping rule principle!
“my second line of argument that there is a reasonable alternative explication of the idea of inference and one which leads to the rejection of sequential analysis. This explication is provided by the likelihood principle—which states that all observations leading to the same likelihood function should lead to the same conclusion.”
I thus went back to the fundamentals (!), namely [freely available] Bernardo’s and Smith’s Section 5.1.4 (reproduced in EJ’s Stopping rule appendix, also citing Cornfield at length), where the likelihood is properly defined by the joint density of the stopping rule τ and the attached sample at their realised values. And failing in the end (and a discussion with Judith)nto spot a missing normalisation constant.
post-Bayes workshop at UCL [15 & 16 May 2025]
Posted in pictures, Statistics, Travel, University life with tags ABC, approximate Bayesian inference, England, generalised Bayesian inference, Great-Britain, Highlands, ICMS, Isle of Skye Brewery, martingale posterior, martingales, PAC, PAC-Bayes, post-Bayes inference, Scotland, Skye, UCL, University College London, workshop on April 4, 2025 by xi'an
University College London (UCL) is organising a workshop on post-Bayes inference and asked to post the announcement (despite my feeling that we have not entered the post-Bayes era!). So here it is:
Over the course of two days, they will host eight invited talks from leaders across the post-Bayesian landscape, spanning from PAC Bayes and generalised Bayes to predictive resampling and martingale posteriors. Alongside these, they will host six contributed talks, and a poster session to ignite discussion and innovation in our growing community. The workshop will complement the post-Bayesian seminar series.
Registration is now open, and they are actively accepting talk and poster submissions! (Deadline for submissions: April 11th, 2025.) Travel support for early career researchers will be available and announced closer to the date. See the website for more information. The workshop will take place in Bentham House, UCL, London.
[As a personal aside, we just learned that our proposal for an approximate(ly) Bayes workshop supported by ICMS (Edinburgh) and set on the magical Isle of Skye had been accepted! To be held in Spring 2026!]
a second course in probability² [book review]
Posted in Books, Kids, Statistics, University life with tags book reviews, Cambridge University Press, central limit theorem, CHANCE, conditional probability, cross validated, cup, dominated convergence, introductory textbooks, Lebesgue integration, Markov chains, martingales, measure theory, minorisation, renewal process, Riemann integration on December 17, 2023 by xi'an
I was sent [by CUP] Ross & Peköz Second Course in Probablity for review. Although it was published in 2003, a second edition has come out this year. I had not looked at the earlier edition hence will not comment on the differences, but rather reflect on my linear reading of the book and my reactions as a potential teacher (even though I have not taught measure theory for decades, being a low priority candidate in an applied math department). As a general perspective, I think it would be deemed as too informal for our 3rd year students in Paris Dauphine.
This indeed is a soft introduction to measure based probability theory. With plenty of relatively basic examples as the requirement on the calculus background of the readers is quite limited. Surprising appearance of an integral in the expectation section before it is ever defined (meaning it is a Riemann integral as confirmed on the next page), but all integrals in the book will be Riemann integrals, with hardly a mention of a more general concept or even of Lebesgue integration (p 16). Which leads to the probability density being defined in terms of the Lebesgue measure (not yet mentioned). Expectation as suprema of step functions which is enough to derive the dominated convergence theorem. And a (insufficiently detailed?) proof that inverting the cdf at a uniform produces a generation from that distribution. Representation that proves most useful for the results of convergence in distribution. Although the choice (p 31) that all rv’s in a sequence are deterministic transforms of the same Uniform may prove challenging for the students (despite mentioning Skohorod’s representation theorem). Concluding the first chapter with an ergodic theorem for stationary and… ergodic sequences, possibly making the result sound circular. Annoyingly (?) a lot of examples involve discrete rvs, the more as we proceed through the chapters. (Hence the unimaginative dice cover.)
Chap 2, the definition of stochastically smaller is missing italics on the term. This chapter relies on the powerful notion of coupling, leading to Le Cam’s theorem and the Stein-Chen method. Declined for Poisson, Geometric, Normal, and Exponential variates, incl. a Central Limit Theorem. Surprising appearance of a conditional distribution and even more of a conditional variate (Theorem 2.11) that I would criticize as sloppy were it to occur within an X validated question!
Chap 3 on martingales with another informal start on conditional expectations using some intuition from the easiest cases, but also a yet undefined notion of conditional distribution. The main application of the notion is the martingale stopping theorem, with mostly discrete illustrations. (The first sentence of the chapter is puzzling, presenting as a generalisation of iid-ness a sequence of rv’s as having each term depending on the previous ones when the joint distribution can always be decomposed this way by a towering argument.)
Chap 4 on probability bounds with a first technique using the importance sampling identity, which includes the Chernoff bound as a special case. While there are principles at work, I am always uncomfortable teaching about these inequalities, as it often relies on a clever trick.
Chap 5 on Markov chains (with Markov deserving of an historical note contrary to Stein or Le Cam, Borel or Cantelli which would have helped my student seeking their names!) but this is solely done on discrete state spaces, without a mention that irreducible transient Markov chains cannot occur on a finite state space. The chapter covers essentials in that context, including Gambler’s ruin, but I’d rather refer to Feller’s (1970) more general coverage and wonder why the authors stuck to the discrete case.
Chap 6 on renewal theory, albeit defined only for crossing renewal times. In the spirit of Meyn & Tweedie (1994), I find renewal times quite useful in establishing Central Limit theorems in non-iid sequences, but here it is only applied to the renewal process itself (with a typo in Proposition 6.7). The chapter however includes an example of forward exact sampling for a Markov chain satisfying a minorisation condition, as well as brief sections on queuing and Poisson processes.
Chap 7 on Brownian motion, no less! With a discrete iterative construction one hopes will conduct to a proper limit as its existence is not formally proven. And which I deem did not require five figures to explain how to randomly move the midpoint of a segment. This short and final chapter proceeds à marche forcée towards a Central Limit theorem for general stationary and ergodic random variables. A bit too much for a 180p book.
[Disclaimer about potential self-plagiarism: this post or an edited version will eventually appear in my Books Review section in CHANCE.]
ABC in Lapland²
Posted in Mountains, pictures, Statistics, University life with tags ABC, ABC gas station, ABC in Lapland, adap'skii, AG:DC workshop, Bayesian bootstrap, Bayesian GANs, Bayesian nonparametrics, causality, clustering, Euler-Marayama discretisation, Finland, FitzHugh-Nagumo process, indirect inference, Kalman filter, label switching, Lapland, Levi, likelihood-free methods, martingales, state space model on March 16, 2023 by xi'an
On the second day of our workshop, Aki Vehtari gave a short talk about his recent works on speed up post processing by importance sampling a simulation of an imprecise version of the likelihood until the desired precision is attained, importance corrected by Pareto smoothing¹⁵. A very interesting foray into the meaning of practical models and the hard constraints on computer precision. Grégoire Clarté (formerly a PhD student of ours at Dauphine) stayed on a similar ground of using sparse GP versions of the likelihood and post processing by VB²³ then stir and repeat!
Riccardo Corradin did model-based clustering when the nonparametric mixture kernel is missing a normalizing constant, using ABC with a Wasserstein distance and an adaptive proposal, with some flavour of ABC-Gibbs (and no issue of label switching since this is clustering). Mixtures of g&k models, yay! Tommaso Rigon reconsidered clustering via a (generalised Bayes à la Bissiri et al.) discrepancy measure rather than a true model, summing over all clusters and observations a discrepancy between said observation and said cluster. Very neat if possibly costly since involving distances to clusters or within clusters. Although she considered post-processing and Bayesian bootstrap, Judith (formerly [?] Dauphine) acknowledged that she somewhat drifted from the theme of the workshop by considering BvM theorems for functionals of unknown functions, with a form of Laplace correction. (Enjoying Lapland so much that I though “Lap” in Judith’s talk was for Lapland rather than Laplace!!!) And applications to causality.
After the (X country skiing) break, Lorenzo Pacchiardi presented his adversarial approach to ABC, differing from Ramesh et al. (2022) by the use of scoring rule minimisation, where unbiased estimators of gradients are available, Ayush Bharti argued for involving experts in selecting the summary statistics, esp. for misspecified models, and Ulpu Remes presented a Jensen-Shanon divergence for selecting models likelihood-freely²², using a test statistic as summary statistic..
Sam Duffield made a case for generalised Bayesian inference in correcting errors in quantum computers, Joshua Bon went back to scoring rules for correcting the ABC approximation, with an importance step, while Trevor Campbell, Iuri Marocco and Hector McKimm nicely concluded the workshop with lightning-fast talks in place of the cancelled poster session. Great workshop, in my most objective opinion, with new directions!