Archive for empirical cdf

Ali–Mikhail–Haq copula, [re]simulated

Posted in Books, R, Statistics with tags , , , , , , , , , on October 13, 2024 by xi'an

When looking for a copula I could simulate from (rather than the Gaussian copula), I found an algorithm for the Ali–Mikhail–Haq copula

C_\theta(u,v) = \frac{uv}{1-\theta(1-u)(1-v)}\quad -1<\theta<1

that was proposed by Kumar (2010) as reproduced above. But the method seriously fails in that the range of (U,V) resulting from the simulation does not even cover the (0,1)² square! Unless I made an R coding mistake (which is always a possibility).

sim=function(T=1e3,h=.5){ 
  o=matrix(runif(2*T),T,2)
  a=1-o[,1];b=1-h*(1+2*a*o[,2])+2*h^2*a^2*o[,2]
  d=1+h*(2-4*a+4*a*o[,2])+h^2*(1-4*a*o[,2] +4*a^2*o[,2])
  o[,2]=1-2*o[,2]*(a*h-1)^2/(b+sqrt(d))
  return(o)}

There is no explanation in the paper as to why this algorithm is (not!) working, besides the inverse cdf argument—with the reference in R copBasic temporarily worrying until I checked the cdf inversion is completely numerical—, but a correct version can be derived from inverting the conditional cdf of one component, V, given the other, U. Namely, since the conditional cdf is given by

F(v|U=u) = \frac{v(1-\theta(1-v))}{(1-\theta(1-u)(1-v))^2}

which leads to a second degree polynomial equation (in v) when solving the equation F(v|U=u) = w.

sim=function(T=1e3,h=.5){
  o=matrix(runif(2*T),T,2)
  v1=o[,1];w=o[,2]
  d=(2*h*v1*w-1-h)^2-4*(1-w)*h*(1-h*w*v1^2)
  o[,2]=-2*h*v1*w+1+h-sqrt(d))/(2*h*(1-h*w*v1^2)
  return(1-o)}

And with a more likely outcome (Xed checked by comparing F(u,v) with its empirical version for several pairs (u,v)):

P(X<Y)

Posted in Books, Kids, Statistics with tags , , , , on July 17, 2023 by xi'an

A simple X-validated question on approximating P(X<Y) from two independent samples of X and Y. Which can be solved by a Monte Carlo approximation to

\mathbb{E}[\mathbb{I}_{X<Y}]

using all pairs of X’s and Y’s if the samples are not too large and a subsample of this set otherwise. However, I am (still) wondering if there is a more efficient way of using both samples, for instance by comparing the empirical cdfs of through using a Mann-Whitney statistic… (ChatGPT version 3.5 first returned a nonsense solution mentioning approximating the above probability when both distributions are equal [!], before acknowledging its error.)

variance of an exponential order statistics

Posted in Books, Kids, pictures, R, Statistics, University life with tags , , , , , , , , , , on November 10, 2016 by xi'an

This afternoon, one of my Monte Carlo students at ENSAE came to me with an exercise from Monte Carlo Statistical Methods that I did not remember having written. And I thus “charged” George Casella with authorship for that exercise!

Exercise 3.3 starts with the usual question (a) about the (Binomial) precision of a tail probability estimator, which is easy to answer by iterating simulation batches. Expressed via the empirical cdf, it is concerned with the vertical variability of this empirical cdf. The second part (b) is more unusual in that the first part is again an evaluation of a tail probability, but then it switches to find the .995 quantile by simulation and produce a precise enough [to three digits] estimate. Which amounts to assess the horizontal variability of this empirical cdf.

As we discussed about this question, my first suggestion was to aim at a value of N, number of Monte Carlo simulations, such that the .995 x N-th spacing had a length of less than one thousandth of the .995 x N-th order statistic. In the case of the Exponential distribution suggested in the exercise, generating order statistics is straightforward, since, as suggested by Devroye, see Section V.3.3, the i-th spacing is an Exponential variate with rate (N-i+1). This is so fast that Devroye suggests simulating Uniform order statistics by inverting Exponential order statistics (p.220)!

However, while still discussing the problem with my student, I came to a better expression of the question, which was to figure out the variance of the .995 x N-th order statistic in the Exponential case. Working with the density of this order statistic however led nowhere useful. A bit later, after Google-ing the problem, I came upon this Stack Exchange solution that made use of the spacing result mentioned above, namely that the expectation and variance of the k-th order statistic are

\mathbb{E}[X_{(k)}]=\sum\limits_{i=N-k+1}^N\frac1i,\qquad \mbox{Var}(X_{(k)})=\sum\limits_{i=N-k+1}^N\frac1{i^2}

which leads to the proper condition on N when imposing the variability constraint.

the random variable that was always less than its mean…

Posted in Books, Kids, R, Statistics with tags , , , , , on May 30, 2016 by xi'an

Although this is far from a paradox when realising why the phenomenon occurs, it took me a few lines to understand why the empirical average of a log-normal sample is apparently a biased estimator of its mean. And why conversely the biased plug-in estimator does not appear to present a bias. To illustrate this “paradox” consider the picture below which compares both estimators of the mean of a log-normal LN(0,σ²) distribution as σ² increases: blue stands for the empirical mean, while gold corresponds to the plug-in estimator exp(σ²/2) when σ² is estimated from the log-sample, as in a normal sample. (The sample is of size 10⁶.) The gold sequence remains around one, while the blue one drifts away towards zero…

The question came on X validated and my first reaction was to doubt an implementation which outcome was so counter-intuitive. But then I thought further about the representation of a log-normal variate as exp(σξ) when ξ is a standard Normal variate. When σ grows large enough, it is near impossible for σξ to be larger than σ². More precisely,

P(X>E[X])=P(σξ>σ²/2)=1-Φ(σ/2)

which can be arbitrarily small.

measuring honesty, with p=.006…

Posted in Books, Kids, pictures, Statistics with tags , , , , , on April 19, 2016 by xi'an


Simon Gächter and Jonathan Schulz recently published a paper in Nature attempting to link intrinsic (individual) honesty with a measure of corruption in the subject home country. Out of more than 2,500 subjects in 23 countries. [I am now reading Nature on a regular basis, thanks to our lab subscribing a coffee room subscription!] Now I may sound naïvely surprised at the methodological contents of the paper and at a publication in Nature but I never read psychology papers, only Andrew’s rants at’em!!!

“The results are consistent with theories of the cultural co-evolution of institutions and values, and show that weak institutions and cultural legacies that generate rule violations not only have direct adverse economic consequences, but might also impair individual intrinsic honesty that is crucial for the smooth functioning of society.”

The experiment behind this article and its rather deep claims is however quite modest: the authors asked people to throw a dice twice without monitoring and rewarded them according to the reported result of the first throw. Being dishonest here means reporting a false result towards a larger monetary gain. This sounds rather artificial and difficult to relate to dishonest behaviours in realistic situations, as I do not see much appeal in cheating for 50 cents or so. Especially since the experiment accounted for a difference in wealth backgrounds, by adapting to the hourly wage in the country (“from $0.7 dollar in Vietnam to $4.2 in the Netherlands“). Furthermore, the subjects of this experiment were undergraduate students in economics departments: depending on the country, this may create a huge bias in terms of social background, as I do not think access to universities is the same in Germany and in Guatemala, say.

“Our expanded scope of societies therefore provides important support and qualifications for the generalizability of these theories—people benchmark their justifiable dishonesty with the extent of dishonesty they see in their societal environment.”

The statistical analysis behind this “important support” is not earth-shattering either. The main argument is based on the empirical cdfs of the gain repartitions per country (in the above graph), with tests that the overall empirical cdf for low corruption countries is higher than the corresponding one for high corruption countries. The comparison of the cumulated or pooled cdf across countries from each group is disputable, in that there is no reason the different countries have the same “honesty” cdf. The groups themselves are built on a rough measure of “prevalence of rule violations”. It is also rather surprising that for both groups the percentage of zero gain answers is “significantly” larger than the expected value of 2.8% if the assumption of “justified dishonesty” holds. In any case, there is no compelling argument as to why students not reporting the value of the first dice would naturally opt for the maximum of the two dices. Hence a certain bemusement at this paper appearing in Nature and even deserving an introductory coverage in the first part of the journal…