Archive for open source

Wikileech

Posted in University life with tags , , , , , , , , on March 17, 2024 by xi'an

Another type of email hassling (or scam?), not asking for payment at this early stage for writing my Wikipedia page for me!, which is going against the rules of the platform:

My name is dah and I work with Wiki-blah, a Wikipedia page creation and management firm. In a quick Google search of your name, I discovered that you do not have a Wikipedia page.

Having a Wikipedia page shows the world that you have made remarkable contributions in your field. Moreover, it catalogues all of your important work in one place, making it easier for the reader to understand your work. Over 73% people searching for an academic on Google visit the academic’s Wikipedia page first and their official university page later. This underscores an important fact: the credibility of Wikipedia surpasses that of your official website. With over 3000 articles declined from Wikipedia everyday, it is not easy to get a Wikipedia page. However, we can make the process very easy for you. We have written Wikipedia pages for over a thousand academics. We can write one for you. For your assurance and safety, we don’t request any upfront payment.

Please send me a message if you would like to find out more about working with us.

And a second one came the day after, in a much more flowery style that omitted any mention of payment:

I trust this message finds you in great health and high spirits.

I’m dah, I am a part of a group of Wikipedia Administrators and Editors, driven by a passion for crafting impactful narratives. With over a decade of experience, our focus lies in assisting individuals and businesses like yours in establishing a lasting presence on Wikipedia, the world’s foremost information platform.

Have you ever considered having your accomplishments and contributions showcased on the influential stage of Wikipedia? I specialize in navigating the intricacies of Wikipedia’s guidelines, ensuring your story meets the stringent standards set by the platform. What sets our approach apart is the inherent resilience of entries created under the supervision of a Wikipedia Administrator – they stand stronger against scrutiny and time. We understand the unique challenges of getting pages published on Wikipedia. We have a proven track record of successfully guiding experts like yourself through the process. By collaborating, we can not only ensure the accurate portrayal of your journey but also secure its place in the annals of Wikipedia’s knowledge repository.

Understanding the complexities of Wikipedia’s communal editing and content standards can be daunting. This is where our expertise shines. As Wikipedia Administrators, we have the ability to guide your entry through the labyrinth of guidelines, maintaining the utmost standards of neutrality and credibility.
If you’re intrigued by the prospect of immortalizing your story on Wikipedia, I’m excited to explore this opportunity further. I’d be happy to share a recent success story or provide a testimonial upon your request. Feel free to reach out with any queries or curiosities you may have. Your achievements deserve a platform that resonates with millions of global readers.

I would love to discuss this opportunity further at your earliest convenience. I look forward to potentially collaborating on this remarkable journey.

banned from the Linux kernel

Posted in Linux, University life with tags , , , , , , , on May 8, 2021 by xi'an

limited shelf validity

Posted in Books, pictures, Statistics, University life with tags , , , , , , , , , , , on December 11, 2019 by xi'an

A great article from Steve Stigler in the new, multi-scaled, and so exciting Harvard Data Science Review magisterially operated by Xiao-Li Meng, on the limitations of old datasets. Illustrated by three famous datasets used by three equally famous statisticians, Quetelet, Bortkiewicz, and Gosset. None of whom were fundamentally interested in the data for their own sake. First, Quetelet’s data was (wrongly) reconstructed and missed the opportunity to beat Galton at discovering correlation. Second, Bortkiewicz went looking (or even cherry-picking!) for these rare events in yearly tables of mortality minutely divided between causes such as military horse kicks. The third dataset is not Guinness‘, but a test between two sleeping pills, operated rather crudely over inmates from a psychiatric institution in Kalamazoo, with further mishandling by Gosset himself. Manipulations that turn the data into dead data, as Steve put it. (And illustrates with the above skull collection picture. As well as warning against attempts at resuscitating dead data into what could be called “zombie data”.)

“Successful resurrection is only slightly more common than in Christian theology.”

His global perspective on dead data is that they should stop being used before extending their (shelf) life, rather than turning into benchmarks recycled over and over as a proof of concept. If only (my two cents) because it leads to calibrate (and choose) methods doing well over these benchmarks. Another example that could have been added to the skulls above is the Galaxy Velocity Dataset that makes frequent appearances in works estimating Gaussian mixtures. Which Radford Neal signaled at the 2001 ICMS workshop on mixture estimation as an inappropriate use of the dataset since astrophysical arguments weighted against a mixture modelling.

“…the role of context in shaping data selection and form—context in temporal, political, and social as well as scientific terms—has been shown to be a powerful and interesting phenomenon.”

The potential for “dead-er” data (my neologism!) increases with the epoch in that the careful sleuth work Steve (and others) conducted about these historical datasets is absolutely impossible with the current massive data sets. Massive and proprietary. And presumably discarded once the associated neural net is designed and sold. Letting the burden of unmasking the potential (or highly probable?) biases to others. Most interestingly, this recoups a “comment” in Nature of 17 October by Sabina Leonelli on the transformation of data from a national treasure to a commodity which “ownership can confer and signal power”. But her call for openness and governance of research data seems as illusory as other attempts to sever the GAFAs from their extra-territorial privileges…

R wins COPSS Award!

Posted in Statistics with tags , , , , , , , , on August 4, 2019 by xi'an

Hadley Wickham from RStudio has won the 2019 COPSS Award, which expresses a rather radical switch from the traditional recipient of this award in that this recognises his many contributions to the R language and in particular to RStudio. The full quote for the nomination is his  “influential work in statistical computing, visualisation, graphics, and data analysis” including “making statistical thinking and computing accessible to a large audience”. With the last part possibly a recognition of the appeal of Open Source… (I was not in Denver for the awards ceremony, having left after the ABC session on Monday morning. Unfortunately, this session only attracted a few souls, due to the competition of twentysome other sessions, including, excusez du peu!, David Dunson’s Medallion Lecture and Michael Lavine’s IOL on the likelihood principle. And Marco Ferreira’s short-course on Bayesian time series. This is the way the joint meeting goes, but it is disappointing to reach so few people.)

Journal of Open Source Software

Posted in Books, R, Statistics, University life with tags , , , , , , , , on October 4, 2016 by xi'an

A week ago, I received a request for refereeing a paper for the Journal of Open Source Software, which I have never seen (or heard of) before. The concept is quite interesting with a scope much broader than statistical computing (as I do not know anyone in the board and no-one there seems affiliated with a Statistics department). Papers are very terse, describing the associated code in one page or two, and the purpose of refereeing is to check the code. (I was asked to evaluate an MCMC R package but declined for lack of time.) Which is a pretty light task if the code is friendly enough to operate right away and provide demos. Best of luck to this endeavour!