Archive for large language models

Nature tidbits [6 Aug 2026]

Posted in Books, pictures, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , on October 7, 2026 by xi'an

The 6 August issue of Nature arrived after my return from Japan, but it took me a while to catch up, the more because its cover story, Bloom service, about fertilising the oceans with iron as a carbon sink did not appeal much to the geoengineering sceptic in me. Nice pun though. Flotsam of interest for me:

  • an article on GPT-4 (and open-weight models) predicting the outcomes of 70 preregistered US survey experiments, 469 effects and 119,330 participants in total. The predicted treatment effects are highly correlated with the observed ones, about as well as pooled human forecasters do, including studies published after the training cutoff. But the predicted effect sizes are systematically too large, which is not a minor detail for a field still recovering from its replication crisis! Plus a survey of 460 social scientists about using LLMs for pilot testing, and a web app to “forecast” your own experiment’s effect before running it. (I wonder what a prior elicited this way would look like.)
  • a paper on membership inference attacks on medical AI models, showing that patients who differ from the majority are the easiest to identify as part of the training data, which I already discussed here. This fits the long-known issue with differential privacy and outliers.
  • a story on the four 2026 Fields Medals, announced at the ICM in Philadelphia, in fields far, far away:
    • Yu Deng (Chicago), for his work on PDEs, including a rigorous derivation of the Boltzmann equation from hard-sphere dynamics (with Zaher Hani and Xiao Ma), wave kinetic equations, and probabilistic approaches to nonlinear Schrödinger equations.
    • John Pardon (Stony Brook), for symplectic topology and enumerative geometry. He started as an undergraduate by answering a question of Gromov on knot distortion.
    • Jacob Tsimerman (Toronto), for extending techniques within arithmetic and complex algebraic geometry and solving conjectures in this field.
    • Hong Wang (NYU and IHES in Bures-sur-Yvette, almost a colleague!), for Fourier restriction, progress on Falconer’s conjecture and, with Joshua Zahl, the resolution of the Kakeya conjecture in three dimensions.
  • a feature on the scientific papers most cited in patents.
  • a news item on a Science paper by Blasi, Hamilton, Gray and Bowern, estimating that humanity may have spoken tens of thousands of languages between 3,000 and 1,000 years ago. The paper suggests a shift away from nomadic life triggered this diversity before it collapsed. I wonder at how deep the statistical linguistics inferring this rich past from the surviving sample goes. In a sense, the finding is not that surprising, given how isolated small prehistoric communities were, and how quickly a language drifts apart without contact. (I would also like to check how much the estimate depends on the assumed birth-and-death rates of languages, since the lost ones leave no trace to calibrate against…)
  • an editorial on the White House’s “new golden age for science” plan. It argues that such a golden age needs open borders and funding for the social sciences, public health and the humanities, not only AI and engineering. Alas, the now usual Trump.2.0 fare, with an AI funding roll-out in the news pages.
  • an emergency call by WHO director-general Tedros Adhanom Ghebreyesus on why the WHO has never been more needed, and news that the first volunteer received an Ebola vaccine three months into the current outbreak.
  • a forecast that the coming El Niño should be the largest (by a mind-blowing margin) on record, pushing 2027 temperatures to new highs. And already having severe impact on tropical weather.
  • a scary comment on the risks of nuclear-powered merchant ships, in the very issue dated on the anniversary of Hiroshima…

Nature squeakbits [26 February 2026]

Posted in Books, pictures, Running, Travel with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on April 16, 2026 by xi'an

In this issue of Nature, uncovering the fundamental (and Ig-Nobel worth) reason why basketball shoes are squeaking, the reason being a shockwave travelling through the sole!, two tribunes against nuclear testing and “must & should” towards a successor to the START treaty… While we should ask & act for a global and if not unilateral nuclear disarmament! This alas coïncides with France announcing the increase of its nuclear arsenal and the extension of its “umbrella” to several EU countries…

And another coverage of the deplorable Trump administration dismantling the biodefense and pandemic preparedness branches of the US NIAID, presumably a late under-the-belt jab at the national institute Anthony Fauci directed for 38 years. Meaning one of the forefronts for pandemic research and vaccine development has been dismantled. Plus the US EPA revoking the 2009 statement that climate change is endangering the US population and removing greenhouse-gas emission rules. In tune with the Drill, baby, drill! motto of the Trump supporters more interested in their short-term profits than in the long-term (no-)future of the country. Paradoxically sitting in the same issue as a comment calling for a policy-making assessment of avoidable climate-change risks. Illustrated by London’s drownin’ below

And a summary of their conclusions from 23 of the 27 members (from 27 countries) of the Scientific Advisory Group for the Origin of Novel Pathogens for the WHO. After 3.5 years of debate on the origin of COVID-19! Four hypotheses are examined and other hoaxes and conspiracies are debunked.

Another political entry about the EU and its Horizon Europe programme (that is funding my ERC Synergy grant) baring (researchers from) Chinese research organisations from applying for its grants in sensitive methodologies, in order to prevent “the undesired transfer of IP”. Following similar and earlier actions in the US with the (newspeak!) China Initiative launched by Trump 1.0 that turned into a witch hunt.

An exciting 228 metres of rock and mud providing a window on the past 23 million year weather. Obtained from the West Antarctica Ice Sheet.

Nature tags & snippets [19 Feb 2026]

Posted in Books, pictures, University life with tags , , , , , , , , , , , , , , , , , , , , , on March 24, 2026 by xi'an

The cover of this 19 Feb edition of Nature is made of an hand pencilled on an Indonesian cave, possibly the oldest piece of art discovered so far. By an ancient precursor of Jean-Michel Basquiat… On more current issues,

A tribune on UKRI cutting whole branches of medical, biological and physical research in the UK, to “focus and do fewer things better¨. Impacting for instance the Medical Research Council (MRC). At least until 2028, with longer range consequences.

And an entry on Epstein, whose files have now contaminated even Nature itself. (When is Andrews MW going to make an appearance there as well?!) The connection (of interest) with science is that the sex offender and financier supported research institutions, incl. the mathematical biologist Martin Nowak, whose work Epstein reviewed (despite a total absence of credentials!). Along another entry on the Chinese Government (MOST) instituting penalties on universities failing to act on scientific misconduct by some of their staff.

And a frightful graph on the exploding number of measles cases in the US, this year, related to the drop in the proportion of vaccinated children. Correlated with the next article on regulation moves by the Trump administration to make firing government scientists much easier. And with the surge in US applicants to the major European grants, namely ERC starting, consolidator and advanced grants and Marie Skłodowska-Curie postdoctoral fellowships.

A book review of Michael Pollan’s A World Appears: A Journey into Consciousness, a New York Times bestseller and more importantly an exploration of the physical nature of consciousness. Followed by a long comment by N. Sanders (Harvard) and B. Schneier (Toronto) on the pernicious effect of high salaries in AI companies on academic research, in particular by negating the scientific worth of collaborating. Plus a regular one-column comment that “Statistical approximation is not general intelligence”..! Albeit with reasonable arguments about the unlikely emergence of an artificial general intelligence.

Last but not least, a Perspective article on the moral competence of large language models. which requires defining first moral competence and moral performances. Involving a long-term acquaintance of mine’s, Kristian Lum (now at DeepMind NY) as one of the authors.

the Harvard and Brown school of computer science

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on October 2, 2025 by xi'an

 “In the late 1980s, LeCun, then a researcher at AT&T Bell Labs, developed a powerful neural network that learned to recognise handwritten zip codes by training on thousands of examples. A parallel development soon unfolded at Harvard and Brown. In 1995, Zhu and a team of researchers there started developing probability-based methods that could learn to recognise patterns and textures (…) and even generate new examples of that pattern. These were not neural networks: members of the “Harvard-Brown school”, as Zhu called his team, cast vision as a problem of statistics and relied on methods such as “Bayesian inference” and “Markov random fields”. The two schools spoke different mathematical languages and had philosophical disagreements. But they shared an underlying logic – that data, rather than hand-coded instructions, could supply the infrastructure for machines to grasp the world and reproduce its patterns – that exists in today’s AI systems such as ChatGPT.” 

watermarking for privacy (with no durian)

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , on February 14, 2025 by xi'an

In Scalable watermarking for identifying large language model outputs, published by Dathathri, See, Ghaisas, et (many) al. in Nature of 23 October 2024, the authors propose an algorithm to (voluntarily) watermark synthetic texts to identify them as such, through a statistical test. Here are a few quotes to relate to the authors’ solution.

“LLMs generate text based on preceding context (…) given a sequence of input text x<t = x1, …, xt−1 consisting of t − 1 tokens from a vocabulary V, the LLM computes the probability distribution pLM(⋅∣x<t) of the next token xt given the preceding text x<t. To generate the full response, xt is sampled from pLM(⋅∣x<t), and the process repeats until either a maximum length is reached or an end-token is generated.

In a watermarking scheme, a sampling algorithm is an algorithm that takes as input a probability distribution p ∈ ΔV and a random seed and returns a token.

Tournament sampling selects a token from the LLM distribution that is likely to score higher under the random watermarking functions (…) Given the selection of tokens xt based on higher g-values, we expect watermarked text generally to score higher under this score than unwatermarked text (…) [It] requires g-values to decide which tokens win each match in the tournament. Intuitively, we want a function that takes a token x ∈ V, a random seed and the layer number ℓ ∈ {1, …, m}, and outputs a g-value gℓ(x, r) that is a pseudorandom sample from some probability distribution fg (the g-value distribution).”

At the Ocean privacy workshop, someone came with the question of trusting (or not) synthetic data and I remembered this article. Suggesting watermarking for said synthetic data by having providers or agents (privately) running a disclosed or registered code that delivers a watermark which provides a strong support to the synthetic data being produced likewise. I remain uncertain this is at realistic.