Archive for natural language processing

punk rock [cover]

Posted in Books, pictures with tags , , , , , , , , , on March 19, 2025 by xi'an

Contextual Integrity for Differential Privacy #4 [23w5106]

Posted in Books, Mountains, pictures, Running, Statistics, Travel, University life, Wines with tags , , , , , , , , , , , , , , , , , , , , , , , , , , on August 5, 2023 by xi'an

Mostly short talks. First talk by Thomas Seinke (Google) on interpreting ε, with a side wondering of mine on the relation between exp(ε) and the uncertainty that comes with Monte Carlo outcome. Which may relate to this 2022 paper by Ruobin Gong. Second talk by Gautam Kamath (U Waterloo) on large language models under privacy with “public” data. Questioning the appropriateness of ML benchmarks in terms of privacy. Third talk by Mark Bun (Boston U) on replicability, privacy and adaptive generalisation in machine learning, with a strange criticism of confidence intervals on the same parameter not intersecting for two independent studies. And proposing high probability replicable algorithms that can be put in duality with differentially private algorithms at the cost of lowering precision and effective sample size. We also had another group discussion on how to reach out about privacy guarantees, which made me realise there were GDPR compliance software available.

In the afternoon session, Shlomi Hod (Boston U) presented a practical case of designing a privacy preserving protocol for the Israeli birth record. With a strong opposition from stakeholders to use synthetic data, due to a semantic drift from synthetic to manipulated to fake, to lying. Wanrong Zhang did not talk about her stunning recent ICML paper but instead of another practical case connected with mobile based Covid case predictions, by adding minimal noise to mobility data. Nidhi Hegde (U Alberta) gave up talking on Thomson sampling with privacy protection, to focus on an ongoing health application for Alberta as more suited for the workshop. And Ria Safavi-Naini (U Calgary) drew a parallel between information theory and DP versus CI.

While the workshop was scheduled till Friday noon, in usual BIRS habits (!), the morning session was cancelled for most people leaving Kelowna in the morning.

more [grass]hope

Posted in Books, Kids, pictures, Statistics with tags , , , , on May 16, 2023 by xi'an

Following my earlier post on the terrible performances of ChatGPT for a grasshopping riddle, Bob Carpenter sent me his interaction with GPT4, using the same entry. Here is the outcome he obtained. The answer is correct, but the reasoning is wrong.

GPT4: This is a classic problem of infinite series and expected values in probability theory.

Given the nature of the grasshopper’s random walk, at each jump, it has an equal probability (1/2) of hopping to the left or to the right. Thus, we expect that, on average, it won’t be moving away from the origin (0) because the positive and negative jumps should cancel out.

However, you are asking for the expected absolute distance from the origin, which can’t be negative by definition. The key to solving this problem is to recognize that the expected value of the absolute value of a random variable is not the same as the absolute value of its expected value.

Since the absolute value function is not linear, we cannot pull the expected value inside of it. Instead, we need to directly compute the expected value of the absolute distance from the origin, which is a sum of an infinite series.

The nth jump’s absolute distance is 1/(2^n), and it has a probability of 1/2 to go either left or right. Therefore, the expected absolute distance from the origin after the nth jump is (1/2) * 1/(2^n).

Adding up these expected distances for all jumps gives us the total expected absolute distance from the origin:

E = (1/2) * (1/2 + 1/4 + 1/8 + 1/16 + …)

The sum inside the parentheses is a geometric series with a = 1/2 and r = 1/2. The sum of an infinite geometric series is a / (1 – r), so the sum of this series is (1/2) / (1 – 1/2) = 1.

Therefore, E = (1/2) * 1 = 1/2

So on average, the grasshopper is expected to be 1/2 unit away from the origin after infinitely many jumps. The expected distance it travels to return to the origin is thus 1/2 unit.