Archive for Sydney Harbour

“too dangerous to release” [a correct prediction]

Posted in Books, Travel with tags , , , , , , , , , , , , on July 9, 2026 by xi'an

Nature of 26 May had a long report on how the Mythos model is deemed to be too dangerous by its conceptor, Anthropic, to be publicly released, but instead shared with a limited number of organisations of their own choice. The argument being that Mythos was able to spot security flaws in every operating system and web browser on the market. But, sooner or later, the AI will prove accessible to a hacker who manages to penetrate one of these organisations, won’t it?! As pointed out by Nature’s writers, the natural question is why should firms be left to decide on who gets access to their software and how dangerous they are. The same writers also muse about governments imposing prohibitions on the firms, including outside their country. Which is exactly what happened a few weeks later on 12 June with the Trump administration “suspend[ing] all access to both Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States” with little factual argumentation. Another illustration of the short-sighted, knee-jerk, inconsistent, policies of that joke of an administration that does not address the real issue.

truly and uniquely pseudo-random?

Posted in Books, pictures, Statistics, Travel, University life with tags , , , , , , , , , , , , , on July 10, 2024 by xi'an

I came across a Web paper entitled The Impact of Google’s Random Number Generator on Accurate Results that I found most puzzling in failing to explain the specific nature of this random generator and that included gems as below

“PRNGs (…) outputs are inherently predictable given a specific seed value, thus deviating from true randomness”

or stating that linear congruential generators suffered from “discernible patterns”, or yet advancing a bonus in using Google’s generator for “leveraging Gaussian distributions”… Referring to a specific and possibly obscure arXival proposing to construct (yet) a PRNG by reinforced learning, with an objective function based on randomness tests reminded me of Marsaglia’s Die Hard. And to “true randomness”. Plus being repetitive and obsequious. Until I found the piece was “written” by Quthor with no typo, an AI “writer” powered by Quick Creator. I should have paid more attention to the level of nonsense…

On the same occurrence, I also read a recent Scientific American paper on randomness tests pseudo-random number generators by Christopher Lutsko. Which does not start that well, since it reproduces the Laplace démon’s argument that a die roll outcome is not random. Also repeating the above pleonasm that PRNGs are not random since they are deterministic. And failing to not mention lava generators! The core of the article is however about recent papers by the author and Niclas Technau of the only examples of sequences that prove passed extremely strong pseudorandom tests. They focus on tests based on gap distribution and pair correlation.

“The [gap distribution] measures the size of the gaps between the points, and the [pair correlation’ measures the clustering of the points—how much they group up or stay apart (…) If these agree with what we would expect from random data, we say that the gap distribution or pair correlation is “Poissonian”.”  Christopher Lutsko

Along with Athanasios Sourmelidis, they proved that exponential sequences {α exp(θ log[n]) mod 1} have Poissonian pair correlation when θ<⅓, for any value of α. And higher Poissonian correlations for θ “smaller and smaller“. (This does not come out of the blue, as the proof relates to Van der Corput’s method of exponential sums, with links to the leading Austrian random generator community.) This is definitely an achievement. However, when θ gets “smaller and smaller“, the terms in the sequence grow more and more slowly, which forces calling for larger and larger values of α to make the generator useful. And the pseudo-randomness test based on Poissonian correlations is only one of many from the toolbox, so the debate seems far from over.

 

 

Approximate Bayesian analysis of (un)conditional copulas [webinar]

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , on September 17, 2020 by xi'an

The Algorithms & Computationally Intensive Inference seminar (access by request) will virtually resume this week in Warwick U on Friday, 18 Sept., at noon (UK time, ie +1GMT) with a talk by (my coauthor and former PhD student) Clara Grazian (now at UNSW), talking about approximate Bayes for copulas:

Many proposals are now available to model complex data, in particular thanks to the recent advances in computational methodologies and algorithms which allow to work with complicated likelihood function in a reasonable amount of time. However, it is, in general, difficult to analyse data characterized by complicated forms of dependence. Copula models have been introduced as probabilistic tools to describe a multivariate random vector via the marginal distributions and a copula function which captures the dependence structure among the vector components, thanks to the Sklar’s theorem, which states that any d-dimensional absolutely continuous density can be uniquely represented as the product of the marginal distributions and the copula function. Major areas of application include econometrics, hydrological engineering, biomedical science, signal processing and finance. Bayesian methods to analyse copula models tend to be computational intensive or to rely on the choice of a particular copula function, in particular because methods of model selection are not yet fully developed in this setting. We will present a general method to estimate some specific quantities of interest of a generic copula by adopting an approximate Bayesian approach based on an approximation of the likelihood function. Our approach is general, in the sense that it could be adapted both to parametric and nonparametric modelling of the marginal distributions and can be generalised in presence of covariates. It also allow to avoid the definition of the copula function. The class of algorithms proposed allows the researcher to model the joint distribution of a random vector in two separate steps: first the marginal distributions and, then, a copula function which captures the dependence structure among the vector components.

 

another Bernoulli factory

Posted in Books, Kids, pictures, R, Statistics with tags , , , , , , on May 18, 2020 by xi'an

A question that came out on X validated is asking for help in figuring out the UMVUE (uniformly minimal variance unbiased estimator) of (1-θ)½ when observing iid Bernoulli B(θ). As it happens, there is no unbiased estimator of this quantity and hence not UMVUE. But there exists a Bernoulli factory producing a coin with probability (1-θ)½ from draws of a coin with probability θ, hence a mean to produce unbiased estimators of this quantity. Although of course UMVUE does not make sense in this sequential framework. While Nacu & Peres (2005) were uncertain there was a Bernoulli factory for θ½, witness their Question #1, Mendo (2018) and Thomas and Blanchet (2018) showed that there does exist a Bernoulli factory solution for θa, 0≤a≤1, with constructive arguments that only require the series expansion of θ½. In my answer to that question, using a straightforward R code, I tested the proposed algorithm, which indeed produces an unbiased estimate of θ½… (Most surprisingly, the question got closed as a “self-study” question, which sounds absurd since it could not occur as an exercise or an exam question, unless the instructor is particularly clueless.)

CRAN does not validate R packages!

Posted in pictures, R, University life with tags , , , , , , , , , , on July 10, 2019 by xi'an

A friend called me the other day for advice on how to submit an R package to CRAN along with a proof his method was mathematically sound. I replied with some items of advice taken from my (limited) experience with submitting packages. And with the remark that CRAN would not validate the mathematical contents of the associated package manual. Nor even the validity of the R code towards delivering the right outcome as stated in the manual. This shocked him quite seriously as he thought having a package accepted by CRAN was a stamp of validation of both the method and the R code. It would be nice of course but would require so much manpower that it seems unrealistic. Some middle ground is to aim at a journal or a peer community validation where both code and methods are vetted. Which happens for instance with the Journal of Computational and Graphical Statistics. Or the Journal of Statistical Software (which should revise its instructions to authors that states “The majority of software published in JSS is written in S, MATLAB, SAS/IML, C++, or Java”. S, really?!)

As for the validity of the latest release of R (currently R-3.6.1 which came out on 2019-07-05, named Action of the Toes!), I figure the bazillion R programs currently running should be able to detect any defect pretty fast, although awareness of the incredible failure of sample() reported in an earlier post took a while to appear.