Wednesday, February 07, 2007

rubinius Interview: rue and IRC Summaries

Recently, rue has been posting summaries of important discussions from the #rubinius irc channel to the rubinius mailing list. This is a great way to catch up the discussions you might have missed.


What made you decide to start writing these irc summaries?

rue: I got back into looking at rubinius a month or so ago after finally getting some free time and did the normal signing up for the ML and idling on IRC routine. After a while I noticed that there is very little public communication despite the yelps of interest from all over Rubyworld and beyond. On IRC, it is usually just the knowledgeable regulars talking.

A good part of it has to do with the subject matter; people are apprehensive about diving in to a compiler project. The catalyst for any thriving newsgroup or ML and thereby the community are questions which yield answers and reciprocal information about the project and with people afraid to ask, the information flow just is not getting established.

So I decided to try to both keep the casual public informed and to try to encourage people to participate whether in discourse or committing code. It is a lot easier to get into rubinius than one would think.

Plus, as you can tell, I like to write.

How do you decide what to keep and what to drop from the summaries?

rue: The summaries are largely the result of personal bias and assumptions I make. Any news about the overall progress and new features and so on are automatically in but aside from that I tend to pick anything that was new or unclear to me or someone else on the channel.

The IRC conversation flow actually makes it fairly easy to highlight the important bits (after some signal to noise reduction practise.)

And some of the stuff in the summaries has nothing to do with IRC except that something there sparked an interest in me to write about it.

Do you have any plans to start putting your summaries into a blog or web format?

rue: Yes, as soon as possible. It is unlikely that I would at this point create any sort of a dedicated structure for these but once some technical problems are resolved, I will add these to my modest journal.

Which threads have you summarized that you think have the most value?

rue: It is hard to ascribe value that way; I like the chats that taught me something and I believe others feel the same way.

Anything that yields follow-ups and gets people involved is the most valuable for the project but it is hard to say which ones do that ahead of time.

Which threads do you remember from before you started summarizing them that you wish you could go back and capture?

rue: Mainly the architectural issues. The project often seems daunting and it will certainly take some time to get one's head fully around it. Anything that sheds light to how and why things are done is excellent material.

Tuesday, February 06, 2007

Parrot/Cardinal Mini-Interview

This is the second week of my break from the serial rubinius interview, and it's time for another piece on parrot and cardinal. Is there any interest in making this into a second (for me, third overall) serial interview? I'd love to do it if there's enough interest.

This week, I'm talking to Kevin Tew (who I interviewed a while ago). In addition to some general information about his Cardinal and Parrot plans, he also gives us his take on VMs supporting multiple languages and register based VMs.


Kevin, now that you're past the current round of heavy lifting involved in school, you've gotten back into Cardinal and Parrot work. Can you tell us a bit about what you've been working on and what's next on your plate?

Kevin: For those who don't know me, I'm a language junkie. I really enjoy working on the low level guts of virtual machines. Lately in Parrot, I have been working on the portion of PIR calling conventions that interface between PIR (parrot high-level assembly language) and C.

Working on Parrot has been an exercise in focus for me. Once you understand the beginnings of how languages, interpreters, and virtual machines work, it is very tempting to start/invent your own language/VM. Working on an existing language/vm project may not be as glamorous as starting your own but for me, choosing to work on Parrot has been very rewarding.

I've spent my most recent Parrot time cage cleaning inter_call.c, which is where Parrot's calling convention and subroutine argument handling code lives.

The cleanup was in preparation for adding PMETHOD support to PMCs. For those that don't know Parrot PMCs are a pseudo C++ objects, implemented in C, upon which much of parrot is build. PMETHODs allow PMC methods, which are written in C, to support the full Parrot Calling Conventions that are available to Parrot Intermediate Representation (PIR, Parrot Assembly). Examples of PIR's enhanced calling capabilities include: optional arguments, named arguments, optional named arguments, slurpy arguments,slurpy named arguments, and array flattening.

What is next on my plate? I'd like to start working on the Parrot Object System. I noticed this past week in Parrot's weekly status meeting that Allison Randal has started working on ParrotBase based on Leo Toetsch CStruct ideas. I better dig in before I get left behind. I started implementing Stevan Little's Class::MOP (Meta Object Protocol) in C PMC's which he of course borrowed from Common Lisp. I soon realized that I wanted full calling convention support in C PMCs so I backtracked and created PMETHODS.

One of the knocks against Parrot is that it's a VM for many languages. It looks like the JVM and rubinius are headed down that road too (to some degree). How does hosting multiple languages help or hurt a VM?

Kevin: Writing a VM for multiple languages is certainly a challenge. The only other comparable VM I can think of that set out to be a VM for multiple languages would be Microsoft's CLR.

Microsoft built the CLR specifically with C# and C++ in mind. VB.NET is also becoming a mature language that is stretching the CLR design. Parrot's multiple language support is a development tactic that has shaped and insured the robustness of Parrot's implemented features. Building a VM, targeted by a spectrum of languages, requires Parrot designers to survey and research how each languages might use or pervert the VM feature in question. While the cost is substantial, I believe it is well worth the price.

Stating multiple language support as a prime object keeps Parrot designers and implementers honest. Corner cutting bits back pretty fast on the Parrot project.

Another knock is that Parrot is register based. Why don't you think this is a problem? (I also asked Dan Sugalski about this.)

Kevin: Stacks are simple to teach and simple to implement. Stack operations, however, are inherently linear. For quite a few years now, all mainstream processors have been superscalar (supporting parallel instruction execution). One of the key roles that caches and register files play in modern processors is to widen the window of instructions which can be analyzed and scheduled for parallel execution. Register based VM architectures remove the overhead of administrative stack operations and map efficiently to the data caches of modern processors.

NOTE: Dan's answer here is better than mine. The above is my opinion, I'm not well versed in this area. I know that stack lovers will say that stacks can map efficiently to register file based processors as well. It was a decision, it's made, life goes on. This issue really gets blown out of proportion. Its not that big of an issue.

Monday, February 05, 2007

Ruby Refactoring Rubicon

I've been watching refactoring in Ruby for quite a while and have long believed that while tools would be nice, Ruby itself made refactoring (and other things) pretty easy. (You can read a bit about that here.) I've also mentioned that this might be the year we get some real refactoring tools. It looks like I might be right.

Martin Fowler wrote an article called The Refactoring Rubicon. He suggested that there are a lot of 'toy' approaches to refactoring built atop regular expressions, but that these will fall over pretty quickly. It's only once you start analyzing the parse tree of a program that you get ability to handle significant refactorings. The simplest refactoring that requires this kind of peeking under the hood is Extract Method, and this is Martin's 'Rubicon'.

Smalltalk crossed the Rubicon long ago. Java crossed it in 2001. And now, Ruby's finally crossed it as well[0]. Jon Tirsén wrote about Ruby crossing the Rubicon. Surprisingly, this advance didn't come out of the Ruby or Smalltalk worlds, it came out of the Java world.

Perhaps, it shouldn't be such a shock. The loudest cries for tools like this have always come from Rubyists with a Java background. Now that JRuby is pushing Ruby deeper into Java-land, it's only natural to see some of the good ideas from the Java community starting to leak into the Ruby community too. It'll be interesting to see what shows up next.

[0] Well, rrb (the Ruby Refactoring Browser) sort of crossed it a couple of years ago, but it was using regular expressions (the regexps that I remembered dealing with were all in the elisp interface to emacs), and wasn't foolproof in its approach (it had a hard time dealing with Extract Method when working with class methods, see bug # 2507) , so I don't think it counts.

Friday, February 02, 2007

Quick rubinius Post

It's stuff like this that keeps reminding me that I should run a tumbleblog. Evan Phoenix has just posted a rubinius style guide ... if you want to hack on rubinius like the rest of the cool kids, you should go read it.

Win a Book and Help rubinius

Want to help the rubinius team? Want another chance to win a book? Well then, this just might be the post for you!.

Every good open source project relies on good names and other cool stuff to keep it running, and rubinius is no exception. We've been trying to come up with a good name for our C API layer for a while now, but we're still stuck with 'RNI' and that's just not good enough.

Now, we're turning to you, intrepid reader — and we're offering a reward. Here's the deal, I'll take submissions for a better name for the C API until midnight (MST) on February 18th (Chinese New Year) then Evan will pick the best name from the batch. All you need to do to submit a book is to enter it in a comment below, you can submit as many names as you like (you're welcome to submit multiple names per comment). The person who's entry is chosen will win their choice of one of these three books:

The (not so) Fine Print:

  • In the event multiple people submit the same name, the first person to have submitted it will be considered the winner.
  • I'm relying on Amazon's free super saver shipping to get the books to you — if it can't ship to where you live, I'm sorry.
  • Pleas avoid names with previous committments, we really don't want to have to go through the hassle of renaming something after Microsoft sends us a Cease &l Desist letter.