“simply start over and build something better”

The post on the shortcomings of R has attracted a huge number of readers and Ross Ihaka has now posted a detailed comment that is fairly pessimistic… Given the radical directions drafted in this comment from the father of R (along with Robert Gentleman), I once again re-post it as a main entry to advertise more broadly its contents. (Obviously, the whole debate is now far beyond my reach! Please comment on the most current post, i.e. this one.)

Since (something like) my name has been taken in vain here, let me chip in.

I’ve been worried for some time that R isn’t going to provide the base that we’re going to need for statistical computation in the future. (It may well be that the future is already upon us.) There are certainly efficiency problems (speed and memory use), but there are more fundamental issues too. Some of these were inherited from Sand some are peculiar to R.

One of the worst problems is scoping. Consider the following little gem.

f =function() {
if (runif(1) > .5)
x = 10
x
}

The x being returned by this function is randomly local or global. There are other examples where variables alternate between local and non-local throughout the body of a function. No sensible language would allow this. It’s ugly and it makes optimisation really difficult. This isn’t the only problem, even weirder things happen  because of interactions between scoping and lazy evaluation.

In light of this, I’ve come to the conclusion that rather than “fixing” R, it would be much more productive to simply start over and build something better. I think the best you could hope for by fixing the efficiency problems in R would be to boost performance by a small multiple, or perhaps as much as an order of magnitude. This probably isn’t enough to justify the effort (Luke Tierney has been working on R compilation for over a decade now).

To try to get an idea of how much speedup is possible, a number of us have been carrying out some experiments to see how much better we could do with something new. Based on prototyping we’ve been doing at Auckland, it looks like it should be straightforward to get two orders of magnitude speedup over R, at least for those computations which are currently bottle-necked. There are a couple of ways to make this happen.

First, scalar computations in R are very slow. This in part because the R interpreter is very slow, but also because there are a no scalar types. By introducing scalars and using compilation it looks like its possible to get a speedup by a factor of several hundred for scalar computations. This is important because it means that many ghastly uses of array operations and the apply functions could be replaced by simple loops. The cost of these improvements is that scope declarations become mandatory and (optional) type declarations are necessary to help the compiler.

As a side-effect of compilation and the use of type-hinting it should be possible to eliminate dispatch overhead for certain (sealed) classes (scalars and arrays in particular). This won’t bring huge benefits across the board, but it will mean that you won’t have to do foreign language calls to get efficiency.

A second big problem is that computations on aggregates (data frames in particular) run at glacial rates. This is entirely down to unnecessary copying because of the call-by-value semantics. Preserving call-by-value semantics while eliminating the extra copying is hard. The best we can probably do is to take a conservative approach. R already tries to avoid copying where it can, but fails in an epic fashion. The alternative is to abandon call-by-value and move to reference semantics. Again, prototyping indicates that several hundredfold speedup is possible (for data frames in particular).

The changes in semantics mentioned above mean that the new language will not be R. However, it won’t be all that far from R and it should be easy to port R code to the new system, perhaps using some form of automatic translation.

If we’re smart about building the new system, it should be possible to make use of multi-cores and parallelism. Adding this to the mix might just make it possible to get a three order-of-magnitude performance boost with just a fraction of the memory that R uses. I think it’s something really worth putting some effort into.

I also think one other change is necessary. The license will need to a better job of protecting work donated to the commons than GPL2 seems to have done. I’m not willing to have any more of my work purloined by the likes of Revolution Analytics, so I’ll be looking for better protection from the license (and being a lot more careful about who I work with).

41 Responses to ““simply start over and build something better””

  1. Regarding Revolution Analytics : They did change some things to R. Nothing stops your from getting the source code and importing their useful changes in free R. After all, they -have to- operate under the GPL2 license as well…

  2. Perhaps a bit out in left field, but I think something Ocaml-esk like a cross platform F# would make a great language underpinning for “new R”. Python is a clean language but is still too slow for complex problems in statistics, simulation, optimisation, and machine learning. Basically python has the same crutch that R has: you need to write outside of python to get performance.

  3. I have been (and remain) a long-time R user and fan. Lately, however, I find myself learning/using Clojure (clojure.org) with Incanter (http://incanter.org/).

    The main reasons for this are (1) memory issues for big data using R, (2) the paradigm of a language deeply rooted in functional programming, (3) concurrency, (4) Java-interoperability as Clojure runs on the JVM, and (5) immediate use of the vast number of incredible Java libraries available (e.g. hadoop, mahout, weka, etc…), to name a few.

    I would be very curious as to how Ross and others view Clojure in the space of the future of statistical programming.

    Many thanks.

  4. John, I fully agree with you. I would very much like to see R moved to a more open license.

    As for the deficiencies in the R language itself – my gosh, just call them bugs, deprecate the old behavior, fix them, and move on. This isn’t the first time a language with significant mindshare has needed some fixing. The challenge is how to do the fixing without losing the mindshare. That takes some careful thought but it’s doable.

    • In principle it would be nice to fix the language deficiencies.

      The only problem might be that that fixes to the languange can (will?) break libraries … so you have to rewrite the libraries … at which point it might be worth to consider to rewrite the libraries in a different programming language so that you don’t have to do any programming language design but only the porting of the libraries to a new language.

      I guess it depends on: Can you improve/fix the language without breaking a significant number of libraries?

    • Mathias- yeah, that has to be the primary concern. In general I don’t think it’s as dire as it seems though. For example, if people are hitting the “sometimes it’s global, sometimes it’s local” bug, then their code is already broken and they’d probably appreciate it being fixed. For other fixes, we’re just talking abstractly, but I doubt rewriting the whole library would usually be required, it’s more likely just a small adjustment.

      For a similar scenario, major releases of perl (which holds the gold-standard-mindshare module repository, CPAN) have deprecated and/or removed features, and a few eggs get broken but in general it’s major forward progress. See http://search.cpan.org/~rgarcia/perl-5.10.0/pod/perl5100delta.pod#Incompatible_Changes .

    • Kevin Wright's avatar
      Kevin Wright Says:

      Amen. I am very surprised by the people calling for more “protection” for code. Are these people currently earning income from code? Would they earn income if R had a more closed license? For better or worse, we live in an age where software is a commodity and the measure of success is not in how much money a person can earn creating code (which seems to be rather little), but in how widely your code is used. Using a restrictive license these days is nearly equivalent to shooting yourself in the foot. Someone else will just come along and offer freer code that will become more widely used.

    • Ken – very good point. I also hope that the language will progress. Perl 5 is an example for well managed progress. Might Perl 6 be an example in the other direction?

      I just tried the equivalent of the function with the scoping problem in incanter:

      (use ‘incanter.stats)
      (defn f [] (do (if (> (sample-uniform 1) 0.5) (def x 10)) x))
      (f)

      Here x always refers to a global symbol, no matter what. (I also think that the code above would not be considered good/idiomatic Clojure code).

      PS: I like to use R and like the user documentation of the available functions. On the other hand I always found it a bit more difficult to get an overview over how the underlying core objects and functions interact to make the larger things click.

  5. Take this with a grain of salt as I have no idea what I’m talking about.

    As I understand it, R in its current incarnation is

    – a language and implementation with its share of warts. E.g., lack of speed, memory management, “dynamic” scoping? (as opposed to lexical); and, in stark contrast,
    – a great library (CRAN).

    I have serious doubts that a “starting over” strategy would work: Noone seems interested to starts such a project, and the community seems to have very different opinions about the direction that such a project should take.

    So what can be achieved with incremental changes? Radford Neal has shown that a lot of speed improvements in R can be made relatively easy. Ross Ihaka otoh points to the underlying difficulties of the language (or its implementation?).

    Would it be possible to change the language of R in iterative steps?
    I’m talking about tightening up, or changing, the semantics of the language, where needed, and deprecate unwanted behaviour, such as the scoping problem. Feature deprecations would be enforced at future milestone releases of R, giving libraries a chance to adapt.

    What I’m suggesting is an evolutionary approach to redesigning R. Sometime in the future it might then make sense to develop a drop-in replacement. Perhaps even develop it concurrently to the evolutionary redesign work.

    That said, when discussing a future reimplementation of R, the LLVM infrastructure seems to be overlooked.

  6. I like the R language and would be glad to see something that’s basically Industrial Strength R, rather than an attempt to use something radically different (say Clojure). Though putting a New R on top of the JVM and enabling tighter integration between Java/Clojure/Scala and New R could be a good thing. (As long as we don’t have to worry about CLASSPATH and as long as we don’t move from packages with most of their logic in opaque C libraries to packages with most of their logic in opaque Java libraries!)

    Another thing I’ve liked about R (in addition to the R community and breadth of packages) is CRAN itself. I’ve installed dozens and dozens of packages and only ever had two that had any complications. Other solutions, based on Python or Java, require more developer-like tools and feel cobbled together. Which is why I’d hesitate to build New R on Python, even if that would enlarge the pool of R developers.

    With R, you might have to worry about a package not being maintained, or having bugs, but you don’t have to worry that you haven’t installed Maven, or haven’t got the right version of some other package — that’s incompatible with other versions of other packages — in order to get this or that to work. With few exceptions, you simply select a (binary) package and load it and run, and you update that package easily (from inside R) and don’t look back or worry about side effects. I feel like I’m able to be a user when I want and a developer when I want and I’m not forced into developer-mode to simply use the tool.

    I’d also agree with John that corporate backing/comfort is important for open source packages. Personally, I think that GNU-style-advocacy only advances itself, not the art.

  7. I think R’s two biggest strengths are it’s data structures (in particular data frames), and it’s libraries. So, in a start-over scenario serious consideration should be given to bridging to another established language to provide the glue. One choice that I have in mind here is numpy/scipy.

    Granted, Python does have short comings – such as the GIL – but it is also well established and refined in many other ways, such as libraries for interfacing into other systems (including GUIs). And, for my money, numpy has demonstrated that python can be quite flexible in terms of language usage.

    From a developers perspective having to learn a whole other language along with it’s idiosyncrasies (R) to perform data mining/analysis tasks, while using another language for most other automation, utility and application tasks (Python) means I am a little less familiar with both than I possibly could be due to time spent with the other.

    Either way, I truly appreciate that R exists in whatever form.

  8. Dear Xi’an,

    Thank you for hosting this fascinating open debate, wonderful job!

    I do have several requests:
    1) It is now unclear to me where to comment to Ross’s comment – here or in the previous article. I will do it there, but if you have a preference – please mention it in the post.

    2) Could you check to see if you might enable threaded commenting on your blog ? (I am not sure your theme supports it. But if it does, you would be able to enable it here:
    http://http://xianblog.wordpress.com/wp-admin/options-discussion.php

  9. I agree with everything Ross says, but I think he is making a big mistake about the license. Revolution Analytics is a net win for the R community, even though they almost certainly choose the wrong horse to bet on. You can download the Revolution R sources here:

    http://www.revolutionanalytics.com/downloads/gpl-sources.php

    If he is complaining the Revolution is releasing some packages as closed products, he’s making a huge mistake thinking that this should be banned. It would effectively kill any use outside academia.

    I actually think that the core of R should be under a very liberal license such as BSD or MIT. Python and numpy’s liberal license allows them to be used without worry within companies and larger applications. Embedding R in larger applications is currently impossible due to license restrictions, which slows contributions to the core of R. Corporate backing has helped many open source projects (see Linux, WebKit, and Java). The goal should be about advancing science and statistical practice, wherever it occurs.

    • I’m in complete agreement with John! It’s not a zero-sum game in that if some third-party company makes money, you lose. It takes a lot of work to make money off of something like R. They also typically assume some risk through endemnification.

      The GPL deters most companies from integrating GPL software. More open licenses promote more users. (I’m speaking in part from experience with LingPipe’s licensing model, which, like AGPL, is even more restrictive than GPL; so much so that some academics can’t abide the restrictions on making the data publicly available.)

      I’d also much rather contribute to something with a BSD/MIT/Apache-type license. It’s more likely I’ll be able to use it myself in the future. The model’s worked well for the Apache web server and the Lucene search engine and the Hadoop map-reduce framework.

      I don’t know Revolution Analytics, but the companies like them in the Apache world like Lucid Imagination have been very helpful in continuing to contribute to the main projects. In some sense, it’s in their best interest not to branch if the main trunk is actively developed.

      Further, by doing serious integration for multiple customers, the companies gain insight that allows them to patch the main open-source versions so that they’re now much more flexible and faster than they were even a couple years ago.

    • It is a good thing if yesterday’s projects are taken over, maintained, and promoted by for-profit organizations. That frees up time of the open source community to develop today’s projects. Just doing patching and maintenance, and continuing to pile things on top of a shaky foundation, are not really attractive activities.

Leave a Reply

Discover more from Xi'an's Og

Subscribe now to keep reading and get access to the full archive.

Continue reading