Archive for the R Category

ECMLE on CRAN

Posted in R, Statistics, University life with tags , , , , , , , , , , on March 27, 2026 by xi'an

x

ChatGPT’ed Monte Carlo exam

Posted in Books, Kids, R, Statistics, University life with tags , , , , , , , , , , on January 22, 2026 by xi'an

This semester I was teaching a graduate course on Monte Carlo methods at Paris Dauphine and I decided to experiment how helpful ChatGPT would prove in writing the final exam. Given my earlier poor impressions, I did not have great expectations and ended up definitely impressed! In total it took me about as long as if I had written the exam by myself, since I went through many iterations, but the outcome was well-suited for my students (or at least for what I expected from my students). The starting point was providing ChatGPT with the articles of Giles on multi-level Monte Carlo and of Jacob et al on unbiased MCMC, and the instruction to turn them into a two-hour exam. Iterations were necessary to break the questions into enough items and to reach the level of mathematical formalism I wanted. Plus add extra questions with R coding. And given the booklet format of the exam, I had to work on the LaTeX formatting (if not on the solution sheet, which spotted a missing assumption in one of my questions). Still a positive experiment I am likely to repeat for the (few) remaining exams I will have to produce!

Approximating evidence via bounded harmonic means (and HPD regions with known volumes)

Posted in pictures, R, Statistics, University life on October 25, 2025 by xi'an

Following a suggestion by Christian Hennig at JSM 2024, I started working with my PhD student Dana Naderi on a detailed assessment of the method we proposed in 2009 with Darren Wraith for evidence approximation. (The method was briefly mentioned in a Physical Review paper and also briefly illustrated in our 2010 San Antonio survey of evidence approximation methods with Jean-Michel Marin.) Well,  it took longer than expected but we eventually completed our paper on the approximation of evidence by bounded harmonic means, exploiting the general identity of Alan Gelfand and Dipak Dey (1994). Following our 2009 idea, the free function in Gelfand & Dey representation is chosen as a Uniform distribution on an HPD region, since this insures boundedness (and hence finite variance) for the resulting estimator. This followed a revival of the method, renamed THAMES, by Metodiev et al. in 2023, where the authors approximate the HPD region with an ellipsoid derived from a Normal distribution centred at the highest of the HPD points, whose covariance matrix is estimated from the posterior sample. With the drawback that this ellipsoid may as well include low probability regions. Our approach (ECMLE, standing for elliptical coverings for marginal likelihood estimation) is aiming at staying within the actual and targeted HPD region by creating non-overlapping ellipsoids from simulations from the posterior and by using a Uniform density on that collection as the reverse importance function. The resulting estimator is unbiased, since the volume of the set is known (when based on a second, independent, sample of simulations from the posterior, a requirement that was not clearly explicated in our earlier survey). In the meanwhile, that is, while we were close to conclude our paper, Metodiev et al. produced a modified version of THAMES, where they truncate the original ellipsoid to intersect with the HPD region of interest, as discussed in an earlier ‘Og entry (with a reply from the authors). While their paper is more focussed on mixture inference and include other aspects on Bayesian inference for mixture, we rewrote ours to include a comparison between the methods, which proves satisfactory, especially in larger dimensions.

Bayes on the Beach 2026 (University of Wollongong, NSW, 9-11 Feb.)

Posted in Kids, Mountains, pictures, R, Running, Statistics, Travel, University life, Wines with tags , , , , , , , , , , , , , , , on September 18, 2025 by xi'an

A modern introduction to probability and statistics [book review]

Posted in Books, R, Statistics, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , on July 12, 2025 by xi'an

In the plane to Bengaluru, I read through the book A modern introduction to probability and statistics, by Graham Upton—whose Measuring Animal Abundance I reviewed for CHANCE a while ago—, which is based on the earlier Understanding Statistics, written jointly with Ian Cook. (Not to be confused with A modern introduction to probability and statistics by Dekking et al.) The subtitle is understanding statistical principles in the computer age. Sorry, in the age of the computer. While the cover is most pleasant (and modern), as noticed by an AF flight attendant, the contents are very very standard and could have been written decades ago since the main concession to “the” computer age is the inclusion of a few R commands at the end of most chapters. There are even a few distribution tables here and there (in case “the” computer is not available). But there is no other connection with computational statistics or statistical computing.

The classicism of the contents and the intended audience mean there is little therein on which to either object or criticise. The mixture of elementary probability and basic statistics in a single textbook always feels awkward to me and I think I would have trouble teaching solely from this material. Apart from the glaring typo on the variance of the sum of two correlated random variables on page 87, missing the factor 2 in front of the covariance, while correct(ed) p97 (and the inevitable “the the” typo spotted once). My main criticisms are on the potential confusion between samples and populations in the early chapters, when some statistics are used as motivational examples, as for instance in a (hidden) Monte Carlo stabilisation to the limiting values (p57), way before the Law of Large Numbers is introduced,, the variable mileage in mathematical rigour (while being uncertain that first year students can handle integrals and derivatives), the textbook examples, and the amount of the book contents spent on descriptive statistics and even more on the “classical” tests, with no critical perspective on using point nulls or p-values. The book concludes with a four page (benevolent) chapter on Bayesian statistics that is superfluous imho, or even counterproductive since my experience with a rushed introduction to Bayesian principles almost always result in a rejection of said principles. Plus, the illustration with the coin tossing is not particularly helpful since Andrew maintains that one can load a die, but cannot bias a coin. (A similar reservation on the half-page 289 coverage on pseudo-random generation and Monte Carlo principles for computing p-values.)

Minor (mostly idiosyncratic) remarks follow: CLT prior to LLN,   n-1 in sample sd, little to no model criticism (ntbcf goodness of fit), missing an opportunity when mentioning the varying probability of a day being a birthday (p31) in contrast with BDA cover story, and another opportunity to cite the 2024 Ig Nobel Prize for coin tossing around the LLN, an unclear definition for random variables( p53) and a potentially confusing introduction of Poisson distributions through a informal reference to Poisson processes (and no reason why the years of accession of the kings of Sussex and England till Guillaume—making a return on p178 with the Domesday Book—in 1066 should follow such a process as suggested in Figure 3.5), a surprising definition of the constant e as the special case of exp(x) when x=1 and its series expansion (p70), omitting proofs on laws of sums of iid rv’s by introducing moment generating functions rather late, another obscure reference to a 16th German treatise on surveying as a precursor of the CLT (p131), a proof for the normalising constant of the Normal density that will most likely escape most first year students, a introduction of the t, F, and χ² distributions with no mention of their respective densities (pp141-147), never defining a joint Normal distribution density, insisting on unbiasedness without noting that maximum likelihood—with a strange motivation that it “makes the next sample of n observations most likely to resemble the data in the current sample (p228)—estimators are almost always biased, an abundance of footnotes that may prove of little interest for the youngest readers.

[Disclaimer about potential self-plagiarism as usual: this post or an edited version will eventually appear in my Books Review section in CHANCE.]