Archive for frogs

Kaw frogs [Wildlife Photographer of the Year]

Posted in Mountains, pictures, Travel with tags , , , , , , , , , , , , , , , , on November 24, 2025 by xi'an

Nature snapshots

Posted in Books with tags , , , , , , , , , , , , , , , , , , , , , , , on September 5, 2024 by xi'an

Some quick breakfast reads from the 11 July issue of Nature (the one with the frog cover!),

  • A tribune on the Canadian example of advocating for graduate and postdoc pay raises, with a success last April. It would prove difficult to achieve in a French academic landscape when postdocs here earn about as much as starting lecturers, whose salary is notoriously low.
  • A news article on the UK elections Labour landslide impact on national scientific landscape, with the former Conservatives’ government Chief Adviser and former Head of Research at GlaxoSmithKline and governmental speaker during the COVID crisis appointed as science minister (how many countries enjoy a science ministry?!) But Labour has shown no inclination to back up (from the former stance) on EU collaborations (no Erasmus!), restrictions on students visa that induced a 40% drop in overseas enrolments, or funding of UK universities (whose finances are in a terrible state). Followed by a call from five UK researchers to “give UK science the overhaul it urgently needs¨.
  • A paper about identifying brain cells attached to a word’s meaning (with an example opposing son and Sun that reminded me of the confusion I had with the Taïwanese movie A Sun!)
  • An another paper on an analysis of January 2020 data collected in the Huanan seafood market in Wuhan, with three virus identified. But inconclusive about the origin of the virus.
  • A work column on the challenges of conservation ecology, with trade-offs between intervention and inaction for endangered species (with the nugget of information that Ecuador gives nature constitutional rights).
  • And a back page on the TIGRR lab [great acronym!] in Melbourne attempt at resurrecting the extinct thyalacine (or Tasmanian tiger) from historical specimens, not mentioning the involvement of a biotech company, Colossal Biosciences, also involved in recreating  mammoths. Which sounds counterproductive beyond the bioengineering feat, esp. with regard to the mammoths, which would be woolly (!) unsuited to the current World (except in some remote corner of Siberia?!) and its climate. When considering how challenging protecting the few remaining African elephants is, imaging free roaming mammoths that would not come to clash with human activities beggars belief.

Nature tidbits

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , , , , , , on September 2, 2024 by xi'an

Going quickly through the four issues of Nature I found in my mailbox when returning from the Pacific Northwest, beyond the great picture of these frogs warming up, and fighting fungal infection, in the 11 July edition, a few (Sunday morn breakfast) quick reads from the 4 July 2024 edition:

  • a “technology and tools” long article on picking the “right” loss function when constructing an algorithmic predictor or another ML tool, although the author remains vague about the “rightness” part (with an incorrect entry for Huber loss). The message is however clearly that tuning the loss function to the problem one wants to address is (obviously) key. With additional advices about properly handling outliers, avoiding overfitting, and correctly modelling noise. In short, run a proper statistical analysis!
  • a call for neuroscientists not to be afraid of studying religions, which should sound like an obviousness, but believers and religious leaders are often so touchy about their faith that running representative scientific studies of the “brain processes associated with religiosity and spirituality” may prove inaccessible. I also find the statement “researchers might be able to get a better handle on what (if any) alterations happen in people’s brains in the rare instances when religious belief turns to radicalized action or sectarian hatred” closer to Brave New World or Clockwork Orange than to the purpose of Nature!
  • the warmest summer ever, ever, ever…
  • a scary story from South Korea where an academic got sentenced to two years in prison for sharing data with Chinese collaborators on “national core technology”, a case reminiscent of similar ones earlier in the US. And where the academic researcher is again seen as the sole culprit instead of their university or the government sharing the responsibility. In France, the imminent threat to turn mathematics and statistics labs into restrictive regime zones (ZRR), with much heavier constraints on visitors and students, and on publication contents, carried by the researchers themselves, pertains from the same tendency. (A personal illustration of the induced administrative absurdities: my datascience lab being next to a biomedical ZRR lab, I am not allowed to use the lift shared by both units, but can nonetheless reach the same spot using the nearby stairs.)

Principles of Applied Statistics

Posted in Books, Statistics, University life with tags , , , , , , , , , , , on February 13, 2012 by xi'an

This book by David Cox and Christl Donnelly, Principles of Applied Statistics, is an extensive coverage of all the necessary steps and precautions one must go through when contemplating applied (i.e. real!) statistics. As the authors write in the very first sentence of the book, “applied statistics is more than data analysis” (p.i); the title could indeed have been “Principled Data Analysis”! Indeed, Principles of Applied Statistics reminded me of how much we (at least I) take “the model” and “the data” for granted when doing statistical analyses, by going through all the pre-data and post-data steps that lead to the “idealized” (p.188) data analysis. The contents of the book are intentionally simple, with hardly any mathematical aspect, but with a clinical attention to exhaustivity and clarity. For instance, even though I would have enjoyed more stress on probabilistic models as the basis for statistical inference, they only appear in the fourth chapter (out of ten) with error in variable models. The painstakingly careful coverage of the myriad of tiny but essential steps involved in a statistical analysis and the highlight of the numerous corresponding pitfalls was certainly illuminating to me.  Just as the book refrains from mathematical digressions (“our emphasis is on the subject-matter, not on the statistical techniques as such p.12), it falls short from engaging into detail and complex data stories. Instead, it uses little grey boxes to convey the pertinent aspects of a given data analysis, referring to a paper for the full story. (I acknowledge this may be frustrating at times, as one would like to read more…) The book reads very nicely and smoothly, and I must acknowledge I read most of it in trains, métros, and planes over the past week. (This remark is not  intended as a criticism against a lack of depth or interest, by all means [and medians]!)

“A general principle, sounding superficial but difficult to implement, is that analyses should be as simple as possible, but not simpler.” (p.9)

To get into more details, Principles of Applied Statistics covers the (most!) purposes of statistical analyses (Chap. 1), design with some special emphasis (Chap. 2-3), which is not surprising given the record of the authors (and “not a moribund art form”, p.51), measurement (Chap. 4), including the special case of latent variables and their role in model formulation, preliminary analysis (Chap. 5) by which the authors mean data screening and graphical pre-analysis, [at last!] models (Chap. 6-7), separated in model formulation [debating the nature of probability] and model choice, the later being  somehow separated from the standard meaning of the term (done in §8.4.5 and §8.4.6), formal [mathematical] inference (Chap. 8), covering in particular testing and multiple testing, interpretation (Chap. 9), i.e. post-processing, and a final epilogue (Chap. 10). The readership of the book is rather broad, from practitioners to students, although both categories do require a good dose of maturity, to teachers, to scientists designing experiments with a statistical mind. It may be deemed too philosophical by some, too allusive by others, but I think it constitutes a magnificent testimony to the depth and to the spectrum of our field.

“Of course, all choices are to some extent provisional.“(p.130)

As a personal aside,  I appreciated the illustration through capture-recapture models (p.36) with a remark of the impact of toe-clipping on frogs, as it reminded me of a similar way of marking lizards when my (then) student Jérôme Dupuis was working on a corresponding capture-recapture dataset in the 90’s. On the opposite, while John Snow‘s story [of using maps to explain the cause of cholera] is alluring, and his map makes for a great cover, I am less convinced it is particularly relevant within this book.

“The word Bayesian, however, became more widely used, sometimes representing a regression to the older usage of flat prior distributions supposedly representing initial ignorance, sometimes meaning models in which the parameters of interest are regarded as random variables and occasionaly meaning little more than that the laws of probability are somewhere invoked.” (p.144)

My main quibble with the book goes, most unsurprisingly!, with the processing of Bayesian analysis found in Principles of Applied Statistics (pp.143-144). Indeed, on the one hand, the method is mostly criticised over those two pages. On the other hand, it is the only method presented with this level of details, including historical background, which seems a bit superfluous for a treatise on applied statistics. The drawbacks mentioned are (p.144)

  • the weight of prior information or modelling as “evidence”;
  • the impact of “indifference or ignorance or reference priors”;
  • whether or not empirical Bayes modelling has been used to construct the prior;
  • whether or not the Bayesian approach is anything more than a “computationally convenient way of obtaining confidence intervals”

The empirical Bayes perspective is the original one found in Robbins (1956) and seems to find grace in the authors’ eyes (“the most satisfactory formulation”, p.156). Contrary to MCMC methods, “a black box in that typically it is unclear which features of the data are driving the conclusions” (p.149)…

“If an issue can be addressed nonparametrically then it will often be better to tackle it parametrically; however, if it cannot be resolved nonparametrically then it is usually dangerous to resolve it parametrically.” (p.96)

Apart from a more philosophical paragraph on the distinction between machine learning and statistical analysis in the final chapter, with the drawback of using neural nets and such as black-box methods (p.185), there is relatively little coverage of non-parametric models, the choice of “parametric formulations” (p.96) being openly chosen. I can somehow understand this perspective for simpler settings, namely that nonparametric models offer little explanation of the production of the data. However, in more complex models, nonparametric components often are a convenient way to evacuate burdensome nuisance parameters…. Again, technical aspects are not the focus of Principles of Applied Statistics so this also explains why it does not dwell intently on nonparametric models.

“A test of meaningfulness of a possible model for a data-generating process is whether it can be used directly to simulate data.” (p.104)

The above remark is quite interesting, especially when accounting for David Cox’ current appreciation of ABC techniques. The impossibility to generate from a posited model as some found in econometrics precludes using ABC, but this does not necessarily mean the model should be excluded as unrealistic…

“The overriding general principle is that there should be a seamless flow between statistical and subject-matter considerations.” (p.188)

As mentioned earlier, the last chapter brings a philosophical conclusion on what is (applied) statistics. It is stresses the need for a careful and principled use of black-box methods so that they preserve a general framework and lead to explicit interpretations.

Error in ABC versus error in model choice

Posted in pictures, Statistics, University life with tags , , , , on March 8, 2011 by xi'an

Following the earlier posts about our lack of confidence in ABC model choice, I got an interesting email from Christopher Drummond, who is a postdoc at University of Idaho, working on an empirical project with the landscape genetics of tailed frogs. Along the lines of the empirical test we advocated at the end of our paper, Chris evaluated the type I error (or the false allocation rate) on a controlled ABC experiment with simulated pseudo-observed data (pods) for validation, and ended up with an large overall error on the order of 10% across four different models, ranging from 5-25% for each.. He further reported that “there is not much improvement of an exponentially decreasing rate of improvement in predictive accuracy as the number of ABC simulations increases” and then extrapolated about the huge [impossibly large] number of ABC  simulations [hence the value of the ABC tolerance] that is required to achieve, say, a 5% error rate. This was a most  interesting extrapolation and we ended up exchanging a few emails around this theme… My main argument in the ensuing discussion was that there is a limiting error rate that presumably is different from zero simply because Bayesian procedures are fallible, just like any other statistical procedure, unless the priors are highly differentiated from one model to the next.

Chris also noticed that calibrating the value of the Bayes factor in terms of the false allocation rate itself rather than an absolute scale like Jeffrey’s might provide some trust about the actual (log10) ABC Bayes factors recovered for the models fit to the actual data he observed, since validation simulations indicated no wrong allocation for values above log(10) BF > 5, versus log10(BF) ~ 8 for the model that best fit the observed data collected from real frogs. Although this sounds like a Bayesian p-value, it illustrates very precisely our suggestion in the conclusion of our paper of turning to empirical measures as such to calibrate the ABC output without overly trusting the ABC approximation of the Bayes factor itself.