Archive for diabetes

Nature’s menu [12 March 2026]

Posted in Kids, pictures with tags , , , , , , , , , , , , , , , , , , , , , , on March 26, 2026 by xi'an


In this issue with a nutrition highlight, some recommendations for healthier options, not particularly surprising:

  • “Morning coffee seems best for heart health“, based on a large longitudinal US study, even though “relationship between coffee consumption and health is unclear”, and especially since this does not impact all-day coffee (and tea?) drinkers.
  • “Go vegan for the gut microbiome“, again based on a relatively large metagenomics study (in the US, the UK, and Italy). Omnivorous get the most diverse microbiomes, but red-meat eaters produce some species linked with IBD and cancers, while vegans host more beneficial bacteria with anti-inflammatory impact. Dairy eaters are also (unsurprisingly) associated with healthier microbiomes. And a connected article in this volume on how changes in the microorganisms in the guts contribute to cognitive decline.
  • “The quest for proteins“, associating the hormone FGF21 as an endocrine signal of protein deprivation, and hence justifying our craving for protein-loaded food. Without concluding at its health consequences.
  • “Sugar rationing reduced diabetes and high blood pressure“, really?! Reminiscing of the post-war (WWII) years in the UK when sugar was rationed. And surveying people born before and after the rationing about their diabetes and hypertension patterns. (Guess what?!)
  • “Ditch the fries, not the mash” as a recommendation to eat potatoes despite the high sugar content of this starchy root (which I very rarely consume, even less in the fried format!). Again based on a huge longitudinal study of 5.2 million people years! The conclusion is still that “replacing total potatoes (…) with whole grains was associated with a lower risk of [type 2 diabetes],”

And a shorter list of recommendations for skin care, away from influencers! Like applying sunscreen, eating a nutrient-dense diet, using a simple, well-balanced moisturizer. Apart from these servings, a continuation of themes met in previous issues

  • an editorial on the three recipients of the 2026 Sony Women in Technology Award with Nature, as part of a series of remarkable women scientists, on the occasion of the International Women’s Day, Xiwen Gong at the University of Michigan, Ellen Roche at the Massachusetts Institute of Technology, and Zhen Xu at the University of Michigan, with a rather un-international location
  • yet another tribune on Epstein!, calling for stricter rules on private funding of research,
  • and yet another on stopping the use of AI in war, which has about as much chance to be heeded as a call to stop the wars (alas!),
  • two further wishful opinion articles calling for action against Trump 2.0, with lots of must and can, but little consideration for the negligence of the rule of law by Agent Orange and his administration…

And an article on how Pokémons inspired future scientists, especially those involved in collecting and classifying.

a journal of the di[y]sruption year

Posted in Books, Kids, Mountains, pictures, Running, Travel, University life, Wines with tags , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , on February 28, 2026 by xi'an


R
ead The Fifth Heart, by Dan Simmons. The same author who wrote the Hyperion Cantos, an impressive creation of a completely alien universe! A winner of both Hugo and Locus Awards. Alas, this book does not belong to the same category. Based on an attractive premise of Henry James meeting Sherlock Holmes, it peters out after a few pages, with no-one at the wheel. The paradox of a famous author meeting a famous character is abandoned after a few pages, Henry James is endlessly wondering about his own writing and position in society, his American friends are all conveniently famous, since they re-enact real characters like John Hay, private secretary of Abraham Lincoln, or the historian Henry Adams. Each new (historical) character comes with a biography complete enough to become a Wikipedia page, just as the geography of each place they visit. The attempt of Henry James at sleuthing is ridiculous and the investigations of both a possible crime (on Adams’s wife, Clover, instead of a suicide) and an anarchist nationwide plot run by… Prof Moriarty are unconvincing. In addition, I find the overall tone of the book highly reactionary and racist, with a fairly long, terrifying, quote from Charles Pearson—unrelated with Karl and Egon—, which could have been intentional to reflect the epoch. But the revisionist approach to the Haymarket massacre adopted by Holmes cannot be seen as such.

Cooked more veg curries and less buckwheat galettes, as my sugar levels went up from the last test! Possibly connected with my intensity training over the period… Still worth trying to cut on high glucose foods! On the home front, I also spent a long morning replacing (and cursing) the (cheap plastic) switch of my vacuum cleaner, despite the guidance of DIY videos where every step seemed so smooth and easy!

Watched the 1995 Western The Quick and the Dead that I found terrible and artificial, despite the collection of famous or soon-famous actors. And a few minutes of two TV Japanese series, SPEC and Silent Truth, not to be continued! As well as skimmed through 96 minutes, a very soapy and un-believable Taiwanese version of Bullet Train Explosion. And listened repeatedly to the guitar adaptation of the Goldberg Variations by Thibaut Garcia & Antoine Morinière, which I first heard on the French Public Radio, France Inter. The pair thought of the project during COVID and had two identical guitars designed and built from the same rosewood tree for the variations. (I have at least six versions of the Goldberg Variations at home, which I discovered via Glenn Gould’s latest interpretation. I remember most fondly driving through Ontario in May 1986, on my way to Ottawa from West Lafayette, and hitting by chance a Canadian radio channel broadcasting the entirety of the 1956 frantic version! As a coïncidence, I found out today—again listening to France Inter and Sophie Marceau—that Thomas Bernhard wrote a novel, Der Untergeher, involving Glenn Gould as one of the characters.)

a first 5k [since last one]

Posted in pictures, Running with tags , , , , , , , , on June 15, 2022 by xi'an

machine-learning harmonic mean

Posted in Books, Statistics with tags , , , , , , on February 25, 2022 by xi'an

In a recent arXival, Jason McEwen propose a resurrection of the “infamous” harmonic mean estimator. In Machine learning assisted Bayesian model comparison: learnt harmonic mean estimator, they propose to aim at the “optimal importance function”. The paper provides a fair coverage of the literature on that topic, incl. our 2009 paper with Darren Wraith (although I do not follow the criticism of using a uniform over an HPD region, esp. since one of the learnt targets is also a uniform over an hypersphere, presumably optimised in terms of the chosen parameterisation).

“…the learnt harmonic mean estimator, a variant of the original estimator that solves its large variance problem. This is achieved by interpreting the harmonic mean estimator as importance sampling and introducing a new target distribution (…) learned to approximate the optimal but inaccessible target, while minimising the variance of the resulting estimator. Since the estimator requires samples of the posterior only it is agnostic to the strategy used to generate posterior samples.”

The method thus builds upon Gelfand and Dey (1994) general proposal that is a form of inverse importance sampling since the numerator [the new target] is free while the denominator is the unnormalised posterior. The optimal target being the complete posterior (since it lead to a null variance), the authors propose to try to approximate this posterior by various means. (Note however that an almost Dirac mass at a value with positive posterior would work as well!, at least in principle…) as the sections on moment approximations sound rather standard (and assume the estimated variances are finite) while the reason for the inclusion of the Bayes factor approximation is rather unclear. However, I am rather skeptical at the proposals made therein towards approximating the posterior distribution, from a Gaussian mixture [for which parameterisation?] to KDEs, or worse ML tools like neural nets [not explored there, which makes one wonder about the title], as the estimands will prove very costly, and suffer from the curse of dimensionality (3 hours for d=2¹⁰…).The Pima Indian women’s diabetes dataset and its quasi-Normal posterior are used as a benchmark, meaning that James and Nicolas did not shout loud enough! And I find surprising that most examples include the original harmonic mean estimator despite its complete lack of trustworthiness.

Leave the Pima Indians alone!

Posted in Books, R, Statistics, University life with tags , , , , , , , , , , , , , , , , on July 15, 2015 by xi'an

“…our findings shall lead to us be critical of certain current practices. Specifically, most papers seem content with comparing some new algorithm with Gibbs sampling, on a few small datasets, such as the well-known Pima Indians diabetes dataset (8 covariates). But we shall see that, for such datasets, approaches that are even more basic than Gibbs sampling are actually hard to beat. In other words, datasets considered in the literature may be too toy-like to be used as a relevant benchmark. On the other hand, if ones considers larger datasets (with say 100 covariates), then not so many approaches seem to remain competitive” (p.1)

Nicolas Chopin and James Ridgway (CREST, Paris) completed and arXived a paper they had “threatened” to publish for a while now, namely why using the Pima Indian R logistic or probit regression benchmark for checking a computational algorithm is not such a great idea! Given that I am definitely guilty of such a sin (in papers not reported in the survey), I was quite eager to read the reasons why! Beyond the debate on the worth of such a benchmark, the paper considers a wider perspective as to how Bayesian computation algorithms should be compared, including the murky waters of CPU time versus designer or programmer time. Which plays against most MCMC sampler.

As a first entry, Nicolas and James point out that the MAP can be derived by standard a Newton-Raphson algorithm when the prior is Gaussian, and even when the prior is Cauchy as it seems most datasets allow for Newton-Raphson convergence. As well as the Hessian. We actually took advantage of this property in our comparison of evidence approximations published in the Festschrift for Jim Berger. Where we also noticed the awesome performances of an importance sampler based on the Gaussian or Laplace approximation. The authors call this proposal their gold standard. Because they also find it hard to beat. They also pursue this approximation to its logical (?) end by proposing an evidence approximation based on the above and Chib’s formula. Two close approximations are provided by INLA for posterior marginals and by a Laplace-EM for a Cauchy prior. Unsurprisingly, the expectation-propagation (EP) approach is also implemented. What EP lacks in theoretical backup, it seems to recover in sheer precision (in the examples analysed in the paper). And unsurprisingly as well the paper includes a randomised quasi-Monte Carlo version of the Gaussian importance sampler. (The authors report that “the improvement brought by RQMC varies strongly across datasets” without elaborating for the reasons behind this variability. They also do not report the CPU time of the IS-QMC, maybe identical to the one for the regular importance sampling.) Maybe more surprising is the absence of a nested sampling version.

pimcisIn the Markov chain Monte Carlo solutions, Nicolas and James compare Gibbs, Metropolis-Hastings, Hamiltonian Monte Carlo, and NUTS. Plus a tempering SMC, All of which are outperformed by importance sampling for small enough datasets. But get back to competing grounds for large enough ones, since importance sampling then fails.

“…let’s all refrain from now on from using datasets and models that are too simple to serve as a reasonable benchmark.” (p.25)

This is a very nice survey on the theme of binary data (more than on the comparison of algorithms in that the authors do not really take into account design and complexity, but resort to MSEs versus CPus). I however do not agree with their overall message to leave the Pima Indians alone. Or at least not for the reason provided therein, namely that faster and more accurate approximations methods are available and cannot be beaten. Benchmarks always have the limitation of “what you get is what you see”, i.e., the output associated with a single dataset that only has that many idiosyncrasies. Plus, the closeness to a perfect normal posterior makes the logistic posterior too regular to pause a real challenge (even though MCMC algorithms are as usual slower than iid sampling). But having faster and more precise resolutions should on the opposite be  cause for cheers, as this provides a reference value, a golden standard, to check against. In a sense, for every Monte Carlo method, there is a much better answer, namely the exact value of the integral or of the optimum! And one is hardly aiming at a more precise inference for the benchmark itself: those Pima Indians [whose actual name is Akimel O’odham] with diabetes involved in the original study are definitely beyond help from statisticians and the model is unlikely to carry out to current populations. When the goal is to compare methods, as in our 2009 paper for Jim Berger’s 60th birthday, what matters is relative speed and relative ease of implementation (besides the obvious convergence to the proper target). In that sense bigger and larger is not always relevant. Unless one tackles really big or really large datasets, for which there is neither benchmark method nor reference value.