
Archive for reproducible research
computo on the go
Posted in Books, R, Statistics, University life with tags binary, Biometrika, Computo, editor in chief, free software, French researchers, github, journal, Julia, Latin, logo, machine learning, open access, open source, Python, Quarto, R, repositories, reproducible research, Rmarkdown, SFDS, Société française de Statistique, Statistics on April 8, 2025 by xi'an
Nice meeting!
Posted in pictures, R, Running, Statistics, Travel, University life with tags Approximate Bayesian computation, Bayesian bootstrap, Bayesian predictive, BIRS, Côte d'Azur, Chennai, Chennai Mathematical Institute, e-values, French Riviera, Gibbs sampling, hypothesis testing, ICSDS 2024, IMS, indistinguishability, Institute of Mathematical Statistics, Neyman-Pearson tests, Nice, open water swimming, p-values, predatory conferences, Promenade des Anglais, reproducible research, Uppsala, w on December 18, 2024 by xi'an
The ICSDS 2024 meeting in Nice is quite impressive and not primarily because it is in Nice under a beautiful December sun. As other (numerous) IMS meetings I attended (since the initial one in Uppsala in 1990!), the program is of high quality and along topics that are currently moving fast or emerging. From the sessions I attended, e-values are strongly represented, although it remains unclear to me why they should constitute a major departure from p-values, as they stick to hypothesis testing, Type I error, power, and the whole paraphernalia of Neyman-Pearson formalism. If I manage to attend a BIRS workshop on the subject next Summer, I may manage to get a better e-derstanding!
The MCMC (only!) session included a presentation by Guanyang Wang that generalised different approximate MCMC schemes into a unified one. And one by Filippo Ascolani on Gibbs beating the competition! I also attended the Bayesian prediction session, where my friends Sonia Petrone and Chris Holmes have presentations on their respective Series B papers. I discussed both on the ‘Og, on 15 March 2023 and 07 November 2022, respectively. This time, I found that both talks had a Bayesian bootstrap flavour, which is not surprising when considering the non-parametric nature of the approach. And they left me wondering at it being protected from overfitting.
My only plenary session was Cynthia Dwork’s on outcome indistinguishability, which, while related to the privacy topics I was topic, remained somewhat obscure as to its purpose. Meaning I have to get through the paper to get a more holistic perspective.
Of course, Nice in Winter is a very nice place, with the waterfront available for running an uninterrupted 15km as we found out with Jérémie Houssineau (at a brisk 4’09” pace I had not planned before starting!) and the sea all for myself (for a dozen minutes before losing digits!). Unfortunately I had to skip the final day due to examinations of the Paris Dauphine MASH master. And miss Stan receiving a student award. But I am looking forward the next iterations of ICSDS. (Not including Copenhagen, Madrid and many many other places in 2025, since ICSDS seemed a most common name for conferences, some presumably predatory! The true location is Sevilla, to keep up with the Mediterranean theme of ICSDS!)
prepaid ABC
Posted in Books, pictures, Statistics, University life with tags ABC, Approximate Bayesian computation, KU Leuven, Leuven, likelihood-free methods, machine learning, neural network, reproducible research, support vector machines, synthetic likelihood on January 16, 2019 by xi'an
Merijn Mestdagha, Stijn Verdoncka, Kristof Meersa, Tim Loossensa, and Francis Tuerlinckx from the KU Leuven, some of whom I met during a visit to its Wallon counterpart Louvain-La-Neuve, proposed and arXived a new likelihood-free approach based on saving simulations on a large scale for future users. Future users interested in the same model. The very same model. This makes the proposal quite puzzling as I have no idea as to when situations with exactly the same experimental conditions, up to the sample size, repeat over and over again. Or even just repeat once. (Some particular settings may accommodate for different sample sizes and the same prepaid database, but others as in genetics clearly do not.) I am sufficiently puzzled to suspect I have missed the message of the paper.
“In various fields, statistical models of interest are analytically intractable. As a result, statistical inference is greatly hampered by computational constraint s. However, given a model, different users with different data are likely to perform similar computations. Computations done by one user are potentially useful for other users with different data sets. We propose a pooling of resources across researchers to capitalize on this. More specifically, we preemptively chart out the entire space of possible model outcomes in a prepaid database. Using advanced interpolation techniques, any individual estimation problem can now be solved on the spot. The prepaid method can easily accommodate different priors as well as constraints on the parameters. We created prepaid databases for three challenging models and demonstrate how they can be distributed through an online parameter estimation service. Our method outperforms state-of-the-art estimation techniques in both speed (with a 23,000 to 100,000-fold speed up) and accuracy, and is able to handle previously quasi inestimable models.”
I foresee potential difficulties with this proposal, like compelling all future users to rely on the same summary statistics, on the same prior distributions (the “representative amount of parameter values”), and requiring a massive storage capacity. Plus furthermore relying at its early stage on the most rudimentary form of an ABC algorithm (although not acknowledged as such), namely the rejection one. When reading the description in the paper, the proposed method indeed selects the parameters (simulated from a prior or a grid) that are producing pseudo-observations that are closest to the actual observations (or their summaries s). The subsample thus constructed is used to derive a (local) non-parametric or machine-learning predictor s=f(θ). From which a point estimator is deduced by minimising in θ a deviance d(s⁰,f(θ)).
The paper does not expand much on the theoretical justifications of the approach (including the appendix that covers a formal situation where the prepaid grid conveniently covers the observed statistics). And thus does not explain on which basis confidence intervals should offer nominal coverage for the prepaid method. Instead, the paper runs comparisons with Simon Wood’s (2010) synthetic likelihood maximisation (Ricker model with three parameters), the rejection ABC algorithm (species dispersion trait model with four parameters), while the Leaky Competing Accumulator (with four parameters as well) seemingly enjoys no alternative. Which is strange since the first step of the prepaid algorithm is an ABC step, but I am unfamiliar with this model. Unsurprisingly, in all these cases, given that the simulation has been done prior to the computing time for the prepaid method and not for either synthetic likelihood or ABC, the former enjoys a massive advantage from the start.
“The prepaid method can be used for a very large number of observations, contrary to the synthetic likelihood or ABC methods. The use of very large simulated data sets allows investigation of large-sample properties of the estimator”
To return to the general proposal and my major reservation or misunderstanding, for different experiments, the (true or pseudo-true) value of the parameter will not be the same, I presume, and hence the region of interest [or grid] will differ. While, again, the computational gain is de facto obvious [since the costly production of the reference table is not repeated], and, to repeat myself, makes the comparison with methods that do require a massive number of simulations from scratch massively in favour of the prepaid option, I do not see a convenient way of recycling these prepaid simulations for another setting, that is, when some experimental factors, sample size or collection, or even just the priors, do differ. Again, I may be missing the point, especially in a specific context like repeated psychological experiments.
While this may have some applications in reproducibility (but maybe not, if the goal is in fact to detect cherry-picking), I see very little use in repeating the same statistical model on different datasets. Even repeating observations will require additional nuisance parameters and possibly perturb the likelihood and/or posterior to large extents.
5 ways to fix statistics?!
Posted in Books, Kids, pictures, Statistics, University life with tags cartoon, falsehood flies and truth comes limping after it, Nature, p-values, poor statistics, predictability, reproducible research, uncertainty on December 4, 2017 by xi'an
In the last issue of Nature (Nov 30), the comment section contains a series of opinions on the reproducibility crisis, by five [groups of] statisticians. Including Blakeley McShane and Andrew Gelman with whom [and others] I wrote a response to the seventy author manifesto. The collection of comments is introduced with the curious sentence
“The problem is not our maths, but ourselves.”
Which I find problematic as (a) the problem is never with the maths, but possibly with the stats!, and (b) the problem stands in inadequate assumptions on the validity of “the” statistical model and on ignoring the resulting epistemic uncertainty. Jeff Leek‘s suggestion to improve the interface with users seems to come short on that level, while David Colquhoun‘s Bayesian balance between p-values and false-positive only address well-specified models. Michèle Nuitjen strikes closer to my perspective by arguing that rigorous rules are unlikely to help, due to the plethora of possible post-data modellings. And Steven Goodman’s putting the blame on the lack of statistical training of scientists (who “only want enough knowledge to run the statistical software that allows them to get their paper out quickly”) is wishful thinking: every scientific study [i.e., the overwhelming majority] involving data cannot involve a statistical expert and every paper involving data analysis cannot be reviewed by a statistical expert. I thus cannot but repeat the conclusion of Blakeley and Andrew:
“A crucial step is to move beyond the alchemy of binary statements about ‘an effect’ or ‘no effect’ with only a P value dividing them. Instead, researchers must accept uncertainty and embrace variation under different circumstances.”
new reproducibility initiative in TOMACS
Posted in Books, Statistics, University life with tags academic journals, ACM, ACM Transactions on Modeling and Computer Simulation, computer, modelling, refereeing, replicating computational results procedure, reproducible research, review, software, TOMACS on April 12, 2016 by xi'an[A quite significant announcement last October from TOMACS that I had missed:]
To improve the reproducibility of modeling and simulation research, TOMACS is pursuing two strategies.
Number one: authors are encouraged to include sufficient information about the core steps of the scientific process leading to the presented research results and to make as many of these steps as transparent as possible, e.g., data, model, experiment settings, incl. methods and configurations, and/or software. Associate editors and reviewers will be asked to assess the paper also with respect to this information. Thus, although not required, submitted manuscripts which provide clear information on how to generate reproducible results, whenever possible, will be considered favorably in the decision process by reviewers and the editors.
Number two: we will form a new replicating computational results activity in modeling and simulation as part of the peer reviewing process (adopting the procedure RCR of ACM TOMS). Authors who are interested in taking part in the RCR activity should announce this in the cover letter. The associate editor and editor in chief will assign a RCR reviewer for this submission. This reviewer will contact the authors and will work together with the authors to replicate the research results presented. Accepted papers that successfully undergo this procedure will be advertised at the TOMACS web page and will be marked with an ACM reproducibility brand. The RCR activity will take place in parallel to the usual reviewing process. The reviewer will write a short report which will be published alongside the original publication. TOMACS also plans to publish short reports about lessons learned from non-successful RCR activities.
[And now the first paper reviewed according to this protocol has been accepted:]
The paper Automatic Moment-Closure Approximation of Spatially Distributed Collective Adaptive Systems is the first paper that took part in the new replicating computational results (RCR) activity of TOMACS. The paper completed successfully the additional reviewing as documented in its RCR report. This reviewing is aimed at ensuring that computational results presented in the paper are replicable. Digital artifacts like software, mechanized proofs, data sets, test suites, or models, are evaluated referring to ease of use, consistency, completeness, and being well documented.