Archive for Poisson regression

Nature [5 Jan issue]

Posted in Books, pictures, University life with tags , , , , , , , , , on February 17, 2023 by xi'an

Nature in its 5 Jan issue has an editorial by Daniël Lakens asking for statistical reviews prior to research being performed and data being collected which sounds like a reasonable idea provided reviewers with proper expertise and dedication can be found, an issue the editorial does not mention. Main focus on sample size that sounds overly simplistic… it contains the following funny (?) jab:

“I do not propose that reviewers debate matters as such as frequentist versus Bayesian philosophies of statistics.”

One could see a connexion with preregistered trials, with the sound argument that hypotheses should be clearly stated prior to getting data.

The issue also contains an open-access paper by WHO and U of Washington researchers (incl. Bayesian John Wakefield) on estimating the number of COVID-19 deaths from excess deaths. With the issue that data is missing for some countries. With a critical commentary from Enrique Acosta on not adjusting for avoided deaths. And apparently (and surprisingly) not accounting for age structure in each country, esp. since regression is involved. The modelling is done via a Poisson count model. And analysed by Bayesian methods. As often I wonder why France doesn’t feature in the picture, except for a mention that the ratio of excess deaths to COVID-19 deaths is less than one, and French Guiana is not on the maps… Unclear issues about highly reliable countries like Germany and Sweden. And splines… Instead of Gaussian processes. No attempt at capture recapture?

And a somewhat puzzling paper [rewarded by the journal cover] on diminishing disruption of scientific papers over time. It is sort of obvious that as the numbers explode novelty and impact diminish. If only because an increasing number of papers never get cited. Based on a single CD index (with a typo in the formula!) Nothing about maths? As noted by the authors in their conclusion the sheer number of disruptive papers had remained essentially constant…

On the use of marginal posteriors in marginal likelihood estimation via importance-sampling

Posted in R, Statistics, University life with tags , , , , , , , , , , , , , on November 20, 2013 by xi'an

Perrakis, Ntzoufras, and Tsionas just arXived a paper on marginal likelihood (evidence) approximation (with the above title). The idea behind the paper is to base importance sampling for the evidence on simulations from the product of the (block) marginal posterior distributions. Those simulations can be directly derived from an MCMC output by randomly permuting the components. The only critical issue is to find good approximations to the marginal posterior densities. This is handled in the paper either by normal approximations or by Rao-Blackwell estimates. the latter being rather costly since one importance weight involves B.L computations, where B is the number of blocks and L the number of samples used in the Rao-Blackwell estimates. The time factor does not seem to be included in the comparison studies run by the authors, although it would seem necessary when comparing scenarii.

After a standard regression example (that did not include Chib’s solution in the comparison), the paper considers  2- and 3-component mixtures. The discussion centres around label switching (of course) and the deficiencies of Chib’s solution against the current method and Neal’s reference. The study does not include averaging Chib’s solution over permutations as in Berkoff et al. (2003) and Marin et al. (2005), an approach that does eliminate the bias. Especially for a small number of components. Instead, the authors stick to the log(k!) correction, despite it being known for being quite unreliable (depending on the amount of overlap between modes). The final example is Diggle et al. (1995) longitudinal Poisson regression with random effects on epileptic patients. The appeal of this model is the unavailability of the integrated likelihood which implies either estimating it by Rao-Blackwellisation or including the 58 latent variables in the analysis.  (There is no comparison with other methods.)

As a side note, among the many references provided by this paper, I did not find trace of Skilling’s nested sampling or of safe harmonic means (as exposed in our own survey on the topic).