Archive for proper scoring rule

watermarking for privacy (with no durian)

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , on February 14, 2025 by xi'an

In Scalable watermarking for identifying large language model outputs, published by Dathathri, See, Ghaisas, et (many) al. in Nature of 23 October 2024, the authors propose an algorithm to (voluntarily) watermark synthetic texts to identify them as such, through a statistical test. Here are a few quotes to relate to the authors’ solution.

“LLMs generate text based on preceding context (…) given a sequence of input text x<t = x1, …, xt−1 consisting of t − 1 tokens from a vocabulary V, the LLM computes the probability distribution pLM(⋅∣x<t) of the next token xt given the preceding text x<t. To generate the full response, xt is sampled from pLM(⋅∣x<t), and the process repeats until either a maximum length is reached or an end-token is generated.

In a watermarking scheme, a sampling algorithm is an algorithm that takes as input a probability distribution p ∈ ΔV and a random seed and returns a token.

Tournament sampling selects a token from the LLM distribution that is likely to score higher under the random watermarking functions (…) Given the selection of tokens xt based on higher g-values, we expect watermarked text generally to score higher under this score than unwatermarked text (…) [It] requires g-values to decide which tokens win each match in the tournament. Intuitively, we want a function that takes a token x ∈ V, a random seed and the layer number ℓ ∈ {1, …, m}, and outputs a g-value gℓ(x, r) that is a pseudorandom sample from some probability distribution fg (the g-value distribution).”

At the Ocean privacy workshop, someone came with the question of trusting (or not) synthetic data and I remembered this article. Suggesting watermarking for said synthetic data by having providers or agents (privately) running a disclosed or registered code that delivers a watermark which provides a strong support to the synthetic data being produced likewise. I remain uncertain this is at realistic.

[strong] foundations of synthetic B’earning

Posted in Books, Statistics, University life with tags , , , , , , , , , , on July 15, 2024 by xi'an

I only recently read the (foundational!) paper on Foundations of Bayesian learnin from synthetic data by Harrison Wilde, Jack Jewson, Sebastian Volmer [all associated with Warwick at some point] and [my long time friend]  Chris Holmes, that merges Bayesian inference with differential privacy constraints thru generalised / Gibbs interface. Recouping with the M-open perspective in order to accommodate the misspecified nature of synthetic data. I like the approach very much in that it intersects a lot with my own views, excepts for following the differential privacy formalism. I however think that further progress could be made by adopting an even more Bayesian position.

Their key messages from that paper are that

  1. learning from synthetic data may prove damaging to your (data) health
  2. robustness unsurprisingly reduces the odds or magnitude of the damage
  3. real data can still be used to some extent

Since the (synthetic) generating model can be a GAN, the privacy requirement is such that noise is “injected” in the input data and in the learning mechanism. This is not discussed in the paper but highly conservative constraints surely make the DGP loose several learning points.

On the side, I also like the alternative of opposing data keeper and learner rather than data owner and adversary. Learning here means taking an optimal B decision about the actual data averaged over the true DGP. With a prior on the distribution of the actual data, while being unable to avoid misspecification in representing the (marginal) synthetic generation model.

Unsurprisingly, the alternative approach is relying on proper scoring rules as Bissiri et al. (2016). Rather than finding the distribution KL closest to the synthetic generative model, robustified by generalised B inference, either via downweighting or via ß-divergence. With a preference for the latter. Interestingly, the authors consider the optimal learning size for the synthetic data. Since bringing in more synthetic data does not mean better performances.

Terry Tao on Bayes… and Trump

Posted in Books, Kids, Statistics, University life with tags , , , , , , , on June 13, 2016 by xi'an

“From the perspective of Bayesian probability, the grade given to a student can then be viewed as a measurement (in logarithmic scale) of how much the posterior probability that the student’s model was correct has improved over the prior probability.” T. Tao, what’s new, June 1

Jean-Michel Marin pointed out to me the recent post of Terry Tao on setting a subjective prior for allocating partial credits to multiple answer questions. (Although I would argue that the main purpose of multiple answer questions is to expedite grading!) The post considers only true-false questionnaires and the case when the student produces a probabilistic assessment of her confidence in the answer. In the format of a probability p for each question. The goal is then to devise a grading principle, f, such that f(p) goes to the right answer and f(1-p) to the wrong answer. This sounds very much like scoring weather forecasters and hence designing proper scoring rules. Which reminds me of the first time I heard a talk about this: it was in Purdue, circa 1988, and Morrie DeGroot gave a talk on scoring forecasters, based on a joint paper he had written with Susie Bayarri. The scoring rule is proper if the expected reward leads to pick p=q when p is the answer given by the student and q her true belief. Terry Tao reaches the well-known conclusion that the grading function f should be f(p)=log²(2p) where log² denotes the base 2 logarithm. One property I was unaware of is that the total expected score writes as N+log²(L) where L is the likelihood associated with the student’s subjective model. (This is the only true Bayesian aspect of the problem.)

An interesting and more Bayesian last question from Terry Tao is about what to do when the probabilities themselves are uncertain. More Bayesian because this is where I would introduce a prior model on this uncertainty, in a hierarchical fashion, in order to estimate the true probabilities. (A non-informative prior makes its way into the comments.) Of course, all this leads to a lot of work given the first incentive of asking multiple choice questions…

One may wonder at the link with scary Donald and there is none! But the next post by Terry Tao is entitled “It ought to be common knowledge that Donald Trump is not fit for the presidency of the United States of America”. And unsurprisingly, as an opinion post, it attracted a large number of non-mathematical comments.