Archive for LLMs

AI in a storm [cover]

Posted in Books, pictures, University life with tags , , , , , , , , , , , , , , , , , , on October 6, 2026 by xi'an

Nature tidbits

Posted in Books, Statistics, University life with tags , , , , , , , , , , , , , , , , , on August 29, 2026 by xi'an

On the 25 June edition, an editorial calling Europe to lead on free and open science, while boosting industrial consequences. (By the way, Japan just joined Horizon Europe! While the UK will rejoin Erasmus in 2027, under the headline “Brexit tore apart European science — now the research rifts are healing“, mentioning the financial and visa unsolved issues.) A news article continuing the investigation of how AI impacts or will impact mathematical research, with a FirstProof test evaluating AIs solving new, research-level, maths problems. (But missing Claude Mythos and Google’s Aletheia.) A “comment” from a Chinese academic that science needs humanities, at a time when universities in the UK are closing some humanities departments. And a discussion of the exceptional discovery of a whale necropolis, 7km deep, with fossil remains dating back to 5M years mixed with recent ones. Plus another exciting 10 page paper about genetic sleuthing on population changes at the collapse of the Roman empire, in the Danube-Isar and Rhein-Main areas of present-day Germany. (The attached picture representing posterior estimates of life expectancies of some individuals, obtained by MCMC.) In the Carreers section, a story from a  geography researcher whose job offer was rescinded at the last moment, a scary thing that is alas not so rare.

Round-robin (with Claude)

Posted in Books, Kids, R, Statistics, Wines with tags , , , , , , , , , , , , on August 15, 2026 by xi'an

A few days ago I had a coffee in Paris with my long-time friend (and former Statistics & Computing editor) Gilles Celeux, and he mentioned me stopping solving and posting maths puzzles like those weekly published by Le Monde. They have indeed vanished with the retirement of the authors, but Gilles added that the arrival of LLMs would have made the exercise moot. I disagreed as (i) the fun of solving the puzzle on my own  has not gone away and (ii) the pedagogical appeal of the puzzle and its resolution remains. As the next Fiddler puzzle arrived in my mailbox, my resolution was put to the test (contrariwise to the previous entry, which did not require massive computations):

The Fiddler League consists of two teams. Over a season, they play each other 162 times. Each team has an equal chance of winning each game, and the results of games are independent. Over the season, on average, how many games would you expect the team with the better record to have won?

I started on the wrong foot with E[X|X≥81] when X is Bin(162,½), equal to 85.77 (either directly or with a Normal approximation), which differs from my second thought, E[max(X,162-X)]=86.07 (either directly or with a Normal approximation), which is larger because of the reflection produced by max. While the first computation was manageable, the second one seemed to involve simulation and I caved in prompting Claude, which provided the answer along with the connection

E[max(X,162−X)]=E[X∣X≥81](1+p81​)−81p81​

After some expansion, the League boasts 30 teams. Over a season, each team plays each other team five times. (Each team plays a total of 145 games.) Again, each team has an equal chance of winning each game, and the results of games are independent. Over the season, on average, how many games would you expect the team with the best record to have won?

The best record is max(Xi) with each of the 30 Xi‘s a sum of 29 Yij and the Yij=5-Yji distributed as Bin(5,½). The Xi‘s are thus Bin (145,½) but dependent. While I could not figure out a closed form answer for the expectation, a direct Monte Carlo resolution is obviously feasible, with Claude (rather than me) running it over 400 million repetitions, but a 30 dimensional Normal approximation exploiting the correlation of 1/29 between the components leads to roughly 85 as the expected value. (Again computed by a Claudicant simulation.)

While the conclusion that the Normal approximation is pretty accurate with so many terms in the Binomial variates is quite unsurprising, Claude saves me coding time without ruining the puzzle altogether. (And Gemini made me aware that the name of the café where Gilles and I regularly meet, L’Écir, is an Auvergne noun for a local, dangerous, mountain blizzard! Thus linking the place to the foundation of the café by Auvergne expatriates…)

arXiv, not AIrXiv!

Posted in Books, University life with tags , , , , , , , on July 18, 2026 by xi'an


As reported in Nature of 28 May, arXiv is now enforcing a ban of researchers (who hardly qualify as “authors”!) submitting articles produced by LLMs with demonstrably insufficient human input, e.g. those with hallucinated (fake) references—easy to detect. The ban is one-year long and, besides, subsequent submissions must first be accepted by a genuine academic journal. Be warned!

“too dangerous to release” [a correct prediction]

Posted in Books, Travel with tags , , , , , , , , , , , , on July 9, 2026 by xi'an

Nature of 26 May had a long report on how the Mythos model is deemed to be too dangerous by its conceptor, Anthropic, to be publicly released, but instead shared with a limited number of organisations of their own choice. The argument being that Mythos was able to spot security flaws in every operating system and web browser on the market. But, sooner or later, the AI will prove accessible to a hacker who manages to penetrate one of these organisations, won’t it?! As pointed out by Nature’s writers, the natural question is why should firms be left to decide on who gets access to their software and how dangerous they are. The same writers also muse about governments imposing prohibitions on the firms, including outside their country. Which is exactly what happened a few weeks later on 12 June with the Trump administration “suspend[ing] all access to both Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States” with little factual argumentation. Another illustration of the short-sighted, knee-jerk, inconsistent, policies of that joke of an administration that does not address the real issue.