Archive for overfitting

Nature tidbits [and garbage out]

Posted in Books, Kids, Mountains, pictures, Travel, University life with tags , , , , , , , , , , , , , , , , , , , , , , , , , on October 25, 2024 by xi'an

Many bits of interest in the 25 July issue of Nature that stayed under a pile of journals for PNW reasons! Even though the cover is rather off-putting…

  • A call by the editors (and computer scientist Cal Newport) for scientists to “stop, drop, and think” (and the paperback version of Newport’s Slow Productivity has a nice cover that looks very much like Moraine Lake in Banff National Parc!)
  • Another call by a Science Po’ sociologist to the (then highly prospective) French government to increase research budget for French universities, which is not particularly original since this is a constant line on all unions’ desiderata, except for the argument that it would make a tuition fees increase more palatable, and to leave more freedom to researchers, with the oft-used argument that “fundamental research may have unexpected implications”. Except that a massive public deficit is going to impose budget squeezes in all spending lines of the Barnier government.
  • Worries of Indian and non-Indian demographers about the continued postponement of India’s national census, with the latest one dating from 2011. With strong repercussions on public policies, especially since the administrative data systems are rarely reliable. The article suggests that more than an estimated 100 million inhabitants are excluded from food subsidies as a consequence. This seems like an opportunity to develop open source alternatives that run the census from indirect observations, bypassing the reluctance of the Modi Government to launch the much delayed census. (Which could have been coupled with the electoral process earlier this year.)
  • An insight on four female “PhD influencers” from the Universities of Pennsylvania, Exeter, Hertfordshire, and Hong Kong (CUHK). Which has a plus side of exposing the life and interests of a PhD student. And a downside of being a social media product, with the many biases that come along posting to the general public.
  • A call for publishing more papers about negative results! Supporting (pre)registration of experiments.
  • A positive book review of The MANIAC of Benjamin Labatut around John von Neumann’s life, analysed by Karl Sigmund
  • A four-page long comment on the importance of preserving scientific US-China relations, despite the growing suspicion in the US that any scientific collaboration with Chinese institutions will benefit Chinese (State) interests. Without falling for a naive approach to such collaborations, and acknowledging the porous boundaries between fundamental and military research, one should acknowledge the great leap forward accomplished by Chinese universities in the past decade that set their researchers at the forefront of science in many fields.
  • A paper related with the cover on model collapse in AIs learning from AI-generated data. Just like MCMC learning a proposal from earlier simulations. A natural consequence of over-fitting.
  • Another paper on the phase (or quantum) transition of the two-dimensional Ising spin glass (model).
  • and the article on predatory conferences I already discussed in August.

no dichotomy between efficiency and interpretability

Posted in Books, Statistics, Travel, University life with tags , , , , , , , , , , , , on December 18, 2019 by xi'an

“…there are actually a lot of applications where people do not try to construct an interpretable model, because they might believe that for a complex data set, an interpretable model could not possibly be as accurate as a black box. Or perhaps they want to preserve the model as proprietary.”

One article I found quite interesting in the second issue of HDSR is “Why are we using black box models in AI when we don’t need to? A lesson from an explainable AI competition” by Cynthia Rudin and Joanna Radin, which describes the setting of a NeurIPS competition last year, the Explainable Machine Learning Challenge, of which I was blissfully unaware. The goal was to construct an operational black box predictor fpr credit scoring and turn it into something interpretable. The authors explain how they built instead a white box predictor (my terms!), namely a linear model, which could not be improved more than marginally by a black box algorithm. (It appears from the references that these authors have a record of analysing black-box models in various setting and demonstrating that they do not always bring more efficiency than interpretable versions.) While this is but one example and even though the authors did not win the challenge (I am unclear why as I did not check the background story, writing on the plane to pre-NeuriPS 2019).

I find this column quite refreshing and worth disseminating, as it challenges the current creed that intractable functions with hundreds of parameters will always do better, if only because they are calibrated within the box and have eventually difficulties to fight over-fitting within (and hence under-fitting outside). This is also a difficulty with common statistical models, but having the ability to construct error evaluations that show how quickly the prediction efficiency deteriorates may prove the more structured and more sparsely parameterised models the winner (of real world competitions).

curve fittings [xkcd]

Posted in Books, Kids with tags , , , , , , on November 4, 2018 by xi'an

JSM 2018 [#4½]

Posted in Statistics, University life with tags , , , , , , , , on August 10, 2018 by xi'an

As I wrote my previous blog entry on JSM2018 before the sessions, I did not have the chance to comment on our mixture session, which I found most interesting!, with new entries on the topic and a great discussion by Bettina Grün. Including the important call for linking weights with the other parameters, as both groups being independent does not make sense when the number of components is uncertain. (Incidentally our paper with Kaniav kamary and Kate Lee does create a dependence.) The talk by Deborah Kunkel was about anchored mixture estimation, a joint work with Mario Peruggia, another arXival that I had missed.

The notion of anchoring found in this paper is to allocate specific observations to specific components. These observations are thus anchored to these components. Among other things, this modification of the sampling model implies a removal of the unidentifiability problem. Hence formally of the label-switching or lack thereof issue. (Although, as Peter Green repeatedly mentioned, visualising the parameter space as a point process eliminates the issue.) This idea is somewhat connected with the constraint Jean Diebolt and I imposed in our 1990 mixture paper, namely that no component would have less than two observations allocated to it, but imposing which ones are which of course reduces drastically the complexity of the model. Another (related) aspect of anchoring is that the observations that are anchored to the components act as parts of the prior model, modifying the initial priors (which can then become improper as in our 1990 paper). The difficulty of the anchoring approach is to find observations to anchor in an unsupervised setting. The paper proceeds by optimising the allocations, which somewhat turns the prior into a data-dependent prior since all observations are used to set the anchors and then used again for the standard Bayesian processing. In that respect, I would rather follow the sequential procedure developed by Nicolas Chopin and Florian Pelgrin, where the number of components grows by steps with the number of observations.

 

JSM 2018 [#1]

Posted in Mountains, Statistics, Travel, University life with tags , , , , , , , , , , on July 30, 2018 by xi'an

As our direct flight from Paris landed in the morning in Vancouver,  we found ourselves in the unusual situation of a few hours to kill before accessing our rental and where else better than a general introduction to deep learning in the first round of sessions at JSM2018?! In my humble opinion, or maybe just because it was past midnight in Paris time!, the talk was pretty uninspiring in missing the natural question of the possible connections between the construction of a prediction function and statistics. Watching improving performances at classifying human faces does not tell much more than creating a massively non-linear function in high dimensions with nicely designed error penalties. Most of the talk droned about neural networks and their fitting by back-propagation and the variations on stochastic gradient descent. Not addressing much rather natural (?) questions about choice of functions at each level, of the number of levels, of the penalty term, or regulariser, and even less the reason why no sparsity is imposed on the structure, despite the humongous number of parameters involved. What came close [but not that close] to sparsity is the notion of dropout, which is a sort of purely automated culling of the nodes, and which was new to me. More like a sort of randomisation that turns the optimisation criterion in an average. Only at the end of the presentation more relevant questions emerged, presenting unsupervised learning as density estimation, the pivot being the generative features of (most) statistical models. And GANs of course. But nonetheless missing an explanation as to why models with massive numbers of parameters can be considered in this setting and not in standard statistics. (One slide about deterministic auto-encoders was somewhat puzzling in that it seemed to repeat the “fiducial mistake”.)