Friday, January 03, 2014

Three thoughts from Ruby Under a Microscope Author, Pat Shaughnessy

I'm reading two great Ruby books right now (reviews will be posted soon): Ruby Under a Microscope
 and  Practical Object-Oriented Design in Ruby. I like them both, for very different reasons.  Today, I'd like to share a little bit about the first.

Pat Shaughnessy (@pat_shaughnessy) has written a great book about how Ruby really works. After getting just a couple of chapters in, I really wanted to pick his brain, and he was kind enough to answer my questions.  Now, I'm sharing his answers with you.

What was the most important thing you learned about Ruby while working on your book?

As I studied Ruby over the past 2 years while working on the book, I’ve been surprised and intrigued by how much of Ruby’s design - and implementation - is based on older computer languages, particularly Smalltalk and Lisp. It’s amazing to me how important the basic computer science research done in the 1960s and 1970s by John McCarthy and his contemporaries really is. It’s almost as if we are all “standing on the shoulders of giants." (Not sure who I'm quoting here…) Learning about these older languages has allowed me to look at Ruby with a new light.

This really hit home for me. I've spent a little (not enough) time looking at Smalltalk and Lisp, and that kind of excursion always makes me think differently about programming (and computing in general).

What do you hope others will take away from reading it?

The main thing I hope readers take away from the book is an understanding and appreciation of the computer science concepts behind how Ruby works. For example: hash tables, garbage collection, closures, stack-based virtual machines and compilers, heap vs. stack memory, LALR parsers… etc. Understanding how Ruby works internally allows me to be more confident while writing Ruby code.

Also, I used a lot of diagrams in Ruby Under a Microscope for two reasons:
  • A picture is worth 1000 words - by using diagrams I was able to explain things better than I could have with just text.
  • But also, and possibly more importantly, I’d like readers to build up a visual model of what Ruby is doing that will pop back into their head the next time they write Ruby code.
Again, this makes me feel better about my side trips into those computer science topics that I enjoy - until they start making my head hurt.

Now that you've successfully published a (really nice looking BTW) Ruby book, what's next for you?

Thanks! Yea, No Starch did a great job with the cover and interior layout. So glad I worked with them.

For right now, I’m just going to be focusing on doing normal Ruby development work. I also hope to learn more about functional languages such as Clojure and Haskell, and I’m also curious to learn more about other new languages such as Go and Rust.

I also want to continue sharing my knowledge in public. I’ll keep blogging, and might try my hand at screencasts and other forms of media. I just finished producing a guest episode for Avdi that will appear in his “Ruby Tapas” series which I’m excited about. Whatever media I end up using, my focus will remain the same: explaining complex topics in a simple way, and touching on the computer science that serves as a foundation to everything we all do today as developers. 

Thursday, May 23, 2013

Post-its and Interviews Part 2

Here's the continuation of Janet and her team getting ready to interview candidates to hire a new team member. (See Part 1 here)
As the team files back in after their break, several people stop in front of the board, looking it over and thinking. Janet calls everyone to the table.

"Ok. We've built a good list here. We've got a couple of tasks to take care of, and maybe a little homework for everyone. Let's start out by talking about our job posting. What should it say? Why don't we break into pairs and see what we can come up with?"

After several minutes the pairs are combined into two teams of four and asked to write a new posting based on both pairs efforts. Several minutes later and the two teams are combined and work on merging their job postings into a final draft.

"Hey, this is a great start, but I think we might want to tailor the message a bit for our different channels," says Cindy, another team member.

"Good idea," responds Janet. "Before we go down that path though, which recruiting channels should we use? Any ideas?"

Several people respond as the lead writes ideas on the white board: the local UX users group, the jobs board at a conference two team members are going to next week, the HR recruiter, and a number of other ideas are all written down. Janet speaks up again.

"I like this list. How should the messages be different for each of them?"

The team dives into the discussion again, coming up with a short list of do's, don'ts, and thoughts about each of the possible recruiting channels. Team members are each assigned a recruiting channel to write a job posting for, and asked to email their efforts to the team tomorrow for approval.

Larry, the senior designer, speaks up, "Ok, if that's out of the way, do we want to build the interviewing team? I really liked the job Susan's been doing leading our reading group. Could she be the primary screening interviewer?"

"I think that would be great! Susan, are you up to it?" Asks Janet.

"Well, I was hoping to be on the main team again, but this sounds like a fun task too." Susan replies.

"Don't worry, we'll look to you for some guidance on what to ask about, and listen for, around books the candidates have read recently." Janet turns to the rest of the team. "Which of our customers should we invite into the interview process?"
Does your team spend enough time before the interview to make sure you'll hire the right people? What do you do to improve your interviewing and hiring process?

Tuesday, May 21, 2013

Post-its and Interviews Part 1

I was in a meeting room I'd not visited before the other day and I saw a great idea on the wall. At a glance, I saw what another team had been doing.  With a little more thought and discussion with a co-worker, I was able to build a more complete picture of their activity, what it could have been about, and what it could lead to.

An entire wall was covered in sticky-notes, each with a short description on it.  There were inscriptions like: "listens", "thoughtful", "understands user perspective", and "can read code". The notes were divided into groups like: "Leadership", "Communications", "Design", and "Information Architecture".

It didn't take much to realize that I was looking at a team's description of what they wanted in a co-worker. It was only a smaller jump to put together the following scenario.

The Web Design Team for Product-X recently had a member jump to another team, and they were preparing to hire her replacement. I imagine Janet, the team lead, gathering everyone in a room and handing out small stacks of sticky notes. Then she would address the team to kick things off.

"You all know what's in our job description. I'd like you to think about that for a moment, then write down some traits you think we need to be looking for. Let's take 5-10 minutes to write them down and put them on the board. Don't be afraid to talk to each other while you're working, we're not keeping any secrets."

I can almost hear the buzz of the team working on this and building on each others ideas. As they wind down, Janet speaks up again.

"Ok, this looks like a good sized list. Let's organize it into groups. If you see duplicates, stack them together and we'll sort them out in a bit."

This is probably a pretty interesting process as clusters of traits are built up, split apart, and moved around. Gradually, the level of activity settles as the team comes to some agreements. Bill, one of the designers might speak up and say something like.
"Hey, I think we're missing some things here in "Leadership" and maybe there in "Design". How about we take five more minutes and flesh these out.  We could work on the duplicates while we're at it."

There's a lot more discussion this time. Some of the supposed duplicates are split off and moved into other groups. Others are recognized as real duplicates and those traits are starred, indicating higher importance. New traits are added to each of the groups. Finally, the energy beings to lag again, and Janet speaks up.

"Ok, this looks like a pretty good start.  I'd like to take a ten minute breather before we come back and get into our next steps."

Everyone files out of the room intent on getting some water, juice, or coffee, and maybe a quick walk outside for some fresh air.

I'll post about the follow-up meeting on Thurssday. Do you think a process like this might help with your next hire? Maybe in writing job descriptions for your team (you do have job descriptions, right?)? Have you tried something like this?

Thursday, November 01, 2012

Reviewing Annual Reviews

Almost everyone seems to hate end of year performance reviews. Done correctly though, they could be celebrations of your accomplishments. What would it take to make them more exciting, more interesting, or at least less painful?

How about this for an annual review?


Maybe we're not going to see videos with voice-over announcers, screaming fans, or a pulsating soundtrack. Surely we can do better than a drab document that briefly mentions the couple of things that we remember from the last year, right?

Here are some thoughts that might make your annual review a more positive experience:
  • Keep records during the year. What have you done? What have people said about it? What did it mean to your org? The more data you have the easier it will be to create a year end review that shines.
  • You might not be the one writing the review, but you can create something to send to your boss to help her remember what you've accomplished this year.
  • Think about how you want to organize your review. You don't have to write things chronologically. Maybe you'd rather pull out some specific themes and follow them?
  • Remember to prioritize your entries by impact too — 2nd, 3rd, 1st is a nice order to help keep the big wins 'top of mind'.
  • Include comments from others. Quoting an email or notes from a on-on-one is a great way to reinforce the value that others see from your contribution.
  • Work from established goals. Both yours and your organization's.
  • Use last year's review and your records from the current year to hold personal retrospectives periodically. Make sure you're on track and excited to move forward.


And, who knows, a soundtrack couldn't hurt, right?

Monday, June 11, 2012

This is a book that I wish was on my son's required reading list.  Not that his code is hard to read (for someone in their first programming class), but that there are all kinds of bad habits that wouldn't need to be broken if he and his classmates spent some time learning what good code looks like before they started to write their own.



The Art of Readable Code from O'Reilly is a quick, easy read with a lot of useful ideas for new programmers.  It weighs in at 180 pages, but there's a lot of well used whitespace and a number of (mostly on topic) comic panels in those pages making it seem shorter.



Part one covers naming, code layout, and writing comments; parts two and three cover the meat of refactoring; and part four discusses testing and gives an example of applying the ideas in the book to a small coding project.



The book's examples are in C++, Python, Java, and Javascript.  I would have appreciated seeing some examples in other languages as well (haskell or scala might be good candidates), especially where that language might obviate or change the advice given.



Truth in advertising note, O'Reilly sent me a free copy of this book to review.


Thursday, June 07, 2012

The Linux Command Line

As a long-time, professional Unix/Linux sysadmin, I spend a lot of time on the commandline. I've grown pretty familiar with it, but I often find that junior teammates don't have the same familiarity. They often grew up in a world of windows and GUIs. That means I spend a lot of time helping them learn the ropes.

When I saw that No Starch Press had published The Linux Command Line, I wanted to take a look and see if it would make a good primer for the new guys on my team. Then, the good folks at No Starch sent me a review copy and I figured I'd better act on it.



This is a decent sized book, weighing in at 430 pages and 36 easily digestible chapters. The writing and layout are pretty good, and there are a ton of examples.

Part One is 10 chapters (~110 pages) that cover the basics: built-ins and simple commands; redirection; command line short-cuts (for bash in emacs-mode); permissions; and processes.

Part Two is 3 chapters that will jumpstart you getting beyond command line use. These talk about: setting and using environment variables and startup files; using vim; and making the command prompt more useful.

Part Three is a bit meatier. This contains 10 chapters covering tools for: package management; networking; working w/ storage media; regular expressions; working with text, and more. A couple of chapters in here seem less than necessary (E.g., Chapter 23 'Compiling Programs') but if you don't need them, you can always skip over them.

Part Four dives into scripting, where the true power of the command line comes out. There are 13 chapters which cover: designing scripts; conditionals; looping; parameters; and more.

My biggest complaint about this book is its bash centricity. It would have been nice to throw in a chapter and some sidebars talking about how other shells differ, or mentioning those parts that only bash provides. Since dash is becoming more frequent on Linux systems, it would seem like a natural comparison to make in the text.

Still, this is a good book on using the Linux Command Line. If you're just getting started with Linux, or you have a friend and/or coworker that needs some help, this is a good place to start.

Thursday, February 16, 2012

The Art of R: interview and mini-review


The Art of R Programming is an approachable guide to the R programming language. While tutorial in nature, it should also serve as a reference.
Author Norman Matloff comes from an academic background, and this shows through in the text. His writing is formal, well organized, and tends toward a pedagogical style. This is not a breezy, conversational book.
Matloff approaches R from a programmer's perspective, rather than a statistician's. This approach shows through in several of the chapters: Ch 9, Object-Oriented Programming; Ch 13, debugging; Ch 14, Performance Enhancement; Ch 15, Interfacing R to other languages; and Ch 16, Parallel R. I do wish he had spoken to using R with Ruby as well as C/C++ and Python. I also would have liked to see a chapter on Functional Programming with R, especially after the teaser in the Introduction.
I asked Norm and an R using friend if they could help me get my head around things a little better, and the following mini-interview is the result.

Almost every language has some kind of math support. Why bother with R? Where does it fit in a programmer's toolkit?
Norm: It's crucial to have matrix support, not necessarily in terms of linear algebra operations but at least having good matrix subsetting capability. MATLAB and the Python extension NumPy have this, but I'm not sure how far they go with it. And since MATLAB is not a free product (in fact very expensive) I'm summarily excluding it anyway. :-)
Second, R has a very rich graphics capability, which really sets it apart from the others. You can see some nice examples (with the underlying R code) in The R Graph Gallery.
Third, R is "statistically correct." It was created by top professional statisticians in industry and academia.
Russel: As something of a polyglot, I find that each language comes with something of an attitude of how problems should be approached. The grammatical structure and keyword vocabulary of each language drives a way of thinking about problems, as well as what sorts of libraries must be created to cover what may be base structures and functions in other languages. R has a particularly rich data representation vocabulary which lends itself very nicely to a data-centric problem solving mindset. While many more general-purpose languages can, with appropriate libraries, deal well with data, R reduces the cognitive load required for working with multidimensional data sets. In my (relatively limited) work with R, I've come to think of R as a domain-specific language that happens to have some general-purpose functionality, while other languages such as Ruby, Python, Perl, etc., are general-purpose languages with many domain-specific libraries.
I really feel drawn to the idea that languages drive approaches to problem solving. It reminds me of the ##PragProg idea of a language of the year. With that in mind, what do you think a dynamic language (Perl, Python, Ruby, etc.) programmer going to find new and different in R? What about a programmer coming from a system programming language (C, C++, etc.)?
Russel There is much in R which is from the "dynamic language" camp you mentioned: dynamically typed variables, an interactive shell, dynamically loaded libraries, etc. These will be pretty quickly noticeable to a C/C++/Java/C# programmer.
The structure and forced-forethought enforced by those languages are part of their value proposition: they force programmers into design paradigms and ways of thinking that scale up well, while dynamic languages, with their looser syntax rules, do not enforce that sort of engineering discipline on the programmer. For highly organized people who think in very structured ways, dynamic languages are "freeing", while less structured thinking programmers can find that the lack of enforced structure puts a lot of onus upon them to be disciplined in their coding as program sizes get larger. For example, a simple flat namespace is great for a small program with a few dozen lines, but namespacing becomes much more important as your programs come to the thousands of lines and dozens of individual functions or components -- especially as programs become the shared workspace of multiple programmers.
I personally use R as a dynamic language, most of the time not even writing programs in it so much as using it in interpreted mode for data analysis and "analysis prototyping." In that sense, R does for data analysis what dynamic languages do for task automation: it allows you to easily play with scenarios and prototype your thinking about data quickly and easily. You can then codify the best of those techniques into a small (or large) program that can automate that work for various data sets.
Similarly, R has a very powerful and interactive help system. Most packages not only have a quickly available set of API and help documents, but sample data sets built right into the library. From a command line, R users can get examples of how to use almost any library, with sample data included specifically for that particular library.
R has some inconsistencies from its history that can make it feel more "old school" in some ways. For example,there are two object models and the older (S3-style) object model is widely used in older libraries. However, it's nowhere near as "bolted-onto" as languages like Perl or C. R has an extremely rich set of libraries easily available via CRAN (a la CPAN), but the flip side of this wealth is that these libraries work in many ways, expecting data in various formats, etc. Again, it's not as spotty as CPAN or the Python Cheese Shop, or even Pear—most packages are quite good— but it can leave some beginners feeling a little lost when they want to accomplish a certain task. That's pretty common in the open source world, of course, but can be an issue.
R's rich first-class data types build a foundation that is nicely added to by the various libraries and simple interactive shell. Enough libraries are written in native code that performance is generally top notch. For my part, I almost always find that the available libraries far exceed my generally limited statistical needs, so I rarely find myself needing to rewrite some particular statistical code. I'm not a statistician, so I find it quite valuable to not have to worry about that aspect of the work I'm doing in any given project. Additionally, the rich libraries generally spur me on to doing a richer analysis of the data than I would if I did not have such a fully-featured tool available.
Norm, in the Introduction of your book, you talk about R as a functional language. I wish there had been a chapter on this. Can you give some examples of what you mean? Russel, do you have any thoughts about R as an FP language?
Russel: Many languages have recognized the value of functional constructs and added at least simple implementations of lambda and map functions, first-class functions and the like . FP is generally considered to be more easily parallelized, and should thus scale better on modern multi-core and CUDA-like systems. This will be quite advantageous in large data processing jobs.
Norm: Every operation in R is a function. For instance, the operation y = x[5]is really the function call y = "["(x,5) Same for + and so on.
This is brought up throughout the book, starting with the vector chapter.
The biggest implication of this, in my opinion, is in performance. One can often speed up a computation by a factor in the hundreds by exploiting the FP nature of R.
What are some of the things you've done with R that show off it's power and/or niche?
Russel R works beautifully for many types of data analysis problems. I recently used R to generate annotated graphs of Bayesian content filter scorings against timestamps, with lowess smooth and regression line and other enhancements, all built into the graphs without additional effort. This was done for all permutations of the 5 variables used in the study which had tens of thousands of data points. I was using this as a script because of my need to regenerate the graphs repeatedly, but before I'd codified that process, I used R in a "tweak and go" sort of way, as R lends itself well to ad hoc data exploration. Adding and removing data attributes, filtering data, generating data models, regressions, etc., are all easy to do in an on-the-fly manner.
Norm: A fun application I've done is R code to analyze the differences and similarities between the various dialects of Chinese. It can be used as a learning aid for those who know one Chinese dialect but not another. This is an example in my book, in the chapter on data frames.

If you're interested in adding R to your arsenal of programming tools, this is a great way to get started.
Truth in posting—No Starch Press sent me a free copy of this book to review.

Friday, April 01, 2011

Protocol Buffers - Brian Palmer's Take

Here's another in my continuing series of Ruby Protocol Buffers posts. This time, I've got an interview with Brian Palmer.(I interviewed Brian about his participation in a 'Programming Death Match' back in 2006.)

While working at Mozy, Brian worked with Protocol Buffers, and now maintains the ruby-protocol-buffer. He left Mozy last September, and is now working at a startup called Instructure now.

I hope you enjoy Brian's take on Protocol Buffers.


How did you get started using Protocol Buffers?

Brian: At Mozy, the back-end storage system is written in C++. When we started to standardize our messaging, we evaluated a few different libraries, like Thrift, but ultimately settled on Protocol Buffers because of their great performance and minimal footprint. Our servers handle terabytes of new data a day, so we became sort of "I/O snobs" I guess. We hated using any library that handled its own I/O in a way that we couldn't pipe everything through our finely tuned event loop and zero-copy data structures. So Protocol Buffers, where the data structure and wire format are standardized but the surrounding protocol is up to you, was a perfect fit.

What are its primary application domains?

Brian: I'd say Protocol Buffers are a better fit than say, JSON, for more performance sensitive code, since the wire format is extremely space efficient but fast to parse and generate. Or if you need a very flexible protocol. For instance, we use protocol buffers for the message header but can just pack the actual file data in a raw byte array in the same message, saving the overhead of having our protocol library parsing all those terabytes of data.

Or if you want a flexible data format that still allows a bit more definition than just an arbitrary map of key/value pairs, like JSON. Requiring that fields be defined up front in the .proto file can be really helpful when trying to coordinate communication between different apps internally, especially with the guaranteed backwards compatibility.

Why did you decide to write a Ruby library?

Brian: Mozy's back-end is C++, but we use Ruby for the integration test suite for that system, along with all the web software. So we found that we really needed Protocol Buffers for Ruby. This was 3+ years ago -- at the time, we looked at the existing ruby-protobuf library and it wasn't at all suitable for our needs.

Initially this was just going to be a small internal tool, part of our testing framework. There wasn't any talk of open sourcing the library until we'd already been using it internally for a couple years. I just looked at ruby-protobuf again when you contacted me, and it looks like it's come a long way in both completeness and performance. Makes me a bit sad that I might have muddied the waters with another competing library when neither is a clear winner, that's unfortunate. Hopefully somebody finds it useful, though.

How much effort did you put into making your library performant? Ruby-ish? What are the trade-offs between the two?

Brian: My main focus was on performance, since we found that ruby's time spent encoding/decoding protocol buffers was actually a bottleneck in running our integration tests. Internally at Mozy we install the library as a debian package, rather than a gem, and this includes the C extension that is currently disabled in the gem packaging, which provides another performance boost, especially on ruby 1.8.7.

I think the library remains pretty ruby-ish though. The only place that doesn't feel very ruby-ish is in the code generation: while in theory you could just write your .pb.rb files by hand without writing a .proto file, it's not very natural in the current implementation. So you'll want to use the code generator. But the runtime API is very natural to use, I think.

Since our main goal was interfacing with an existing C++ infrastructure that uses .proto files extensively, this was a natural trade-off for us to make. It wouldn't take much effort to make a more ruby-ish DSL for the modules, though.

With three competing implementations out there, how do you see the playing field shaking out?

Brian: To be honest, I haven't looked closely at the other ruby offerings in a couple years. But I can say that while generally I'm a big fan of choice, I'm not sure it really makes sense to have three ruby libraries for something as simple as protocol buffers — one library could probably easily be made to serve everybody's purposes. So in that sense, I hope a clear winner is established, if just to avoid fragmentation of effort.

What would you like to see happen in terms of documentation for ruby-protocol-buffers?

Brian: The best documentation for Protocol Buffers in general is definitely going to remain on Google's site. And my online library documentation covers how to use the runtime API pretty well. But it'd be nice to have more explanation of how to get up and running with ruby-protocol-buffers.

Thursday, March 31, 2011

Protocol Buffers - BJ Neilsen's Take

BJ Neilsen (@localshred or at github) is a member of my local Ruby Brigade, and he's hacking with/on Protocol Buffers with Ruby — oh, and he's a fan of Real Salt Lake too.

He works for a Provo, Utah based startup MoneyDesktop. Where he helped them transition away from a less-than-desirable PHP solution to Rails. They now enjoy an entirely new service-architecture driven by Ruby (and Protobuf). When not working with Ruby, he runs OneSimpleGoal and plays around with iOS and Objective-C.

To get another take on Protocol Buffers, I asked BJ to join me for a quick interview. Enjoy!


How did you get started using Protocol Buffers?

BJ: At the beginning of 2010 I was hired by a startup in Provo to help build out their product offering. The entire application was written in Java, but for the piece I was to be in charge of I was given free reign to choose a platform. Of course I chose Ruby, but it soon became apparent that we needed a solid way to get data from one application to the other.

This need launched a refactor to a more service-oriented approach. Different solutions were researched for dealing with data interchange such as Thrift and the like, but we ended up choosing Protobuf for its simplicity, pedigree, and multi-platform support. No XML, no WSDL, just simple definitions compiled to the language of your choice. Defining a Data Structure and API with one declarative language, and then being able to build the client and server implementations in two different languages was a huge win. We created a Socket-based RPC server on the Java side, and called the endpoints from Ruby. It was very simple.

I'm now with a new company and the new team was very receptive to the idea of a Protobuf Service ecosystem for our service-oriented application. It is currently the primary method of internal data interchange between multiple service applications. At the time of writing, we have over 20 different proto definition files, 63 separate defined data types (including Enums), 15 independent service classes implementing a total of 32 service endpoints.

What do you see as the strengths of the Protocol Buffers data format?

BJ: One of the greatest strengths of Protobuf is its clear data definitions. Open up any .proto file and it's not hard to deduce the structure of the represented Data Types. Defining Service endpoints is similarly simple, meaning all of the ambiguity of WIki-based (or similar) API documentation is immediately eliminated. Clarity is such a key when building a large system with a team of any size. Being able to clearly understand how and what data is transferred within the system is absolutely key, especially when you hire beyond your core development team and need to get people contributing quickly.

I've already mentioned the power we gained from being able to tie together a Service architecture with multiple languages in a unified API. The Protobuf project officially supports Java, C++, and Python implementations for the definitions compiler and data serialization code, but they have a ton of third party code listed for many other languages like Objective-C and JavaScript (with support in Node.js as well).

Which Protocol Buffers implementation are you using? How did you end up choosing it?

BJ: The only Ruby project listed on Protobuf's "Third Party" page (at the time) was Mack's Ruby-Protobuf. This was a great start as the compiler was built in YACC. However, once I started integrating the API into our Ruby application, it became clear that the RPC side had been half-baked and just sort of thrown out into the wild. Files were compiled and stubbed in the wrong places, meaning that if I added any code to the stubbed client or server files, subsequent compiles would overwrite my changes. Not good.

By that time we were full-steam ahead on the Protobuf implementation in the other services, so I basically had to go in and rewrite the compiler code generation for each of the services, as well as a complete rewrite the entire RPC backend to become compatible with the Protobuf SocketRPC library written for Java. Since that first rewrite at the early part of 2010, I've since done another rewrite (late 2010) to use EventMachine as the RPC backend and I can tell you its lightyears faster, and the DSL is much sexier also, looking much more like an AJAX request with callbacks than a standard socket connection with byte-reading hell. You can get that code on my github fork on the compatability-0.4.0 branch.

What are your plans for you fork of Mack's ruby-protobuf? Will it get wrapped into his distribution or will you go all the way, rename it, and start publishing it as a gem?

BJ: Fantastic question. Currently I've packaged the gem internally for our SOA ecosystem to get around the problem of getting it into a full release with the original code. I've embarked in merge-hell attempting to get my code to work with theirs several times now and each time it just feels like it's not worth it. I've yet to have contact with the original developers (I'm fairly sure they live in Japan) and so I'm not entirely sure they'd accept any patches I'd send anyways.

I've also toyed with the idea that since I've changed a significant chunk of the original code I could just make it my own gem with some witty name (and a reference to the original). The only thing that has kept me from that path is that a) I'd prefer not to insult the original developers, and b) I'm a bit ashamed that there aren't very many tests backing up the RPC backend (the major piece that I wrote from scratch).

Each day we have thousands of successful RPC calls with a virtually non-existent error rate running through the EventMachine RPC code written into this gem, so it has certainly been battle tested in a heavily used production system. Unfortunately it just doesn't have that warm fuzzy feeling (for those who haven't used it yet) that you get when you have 200 green tests behind each class. However, patches with tests are certainly welcome :).

Anyone can pull from my fork on the compatibility-0.4.0 branch (essentially my "master" I build the gem from) and build their own gem if they wish. The current version in my fork is 0.4.0.8. I'd be happy to provide any answers to questions that may arise, and I may even be available to consult with anyone on how to implement Protobuf into your current system.

You gave a presentation on Protocol Buffers at uv.rb. How was it received? Do you see more people starting to use this data format?

BJ: To be honest, I'm not sure my presentation went the way I'd hoped, certainly not well enough to highlight many of the benefits and reasons for using Probotuf. I spent too much time showing the "How" instead of the "Why". I think many people left the meeting intrigued but it was also marred by a drawn-out rant by a few of the developers that were present, debating whether or not it was more prudent to use REST/JSON than a more declarative format like Protobuf.

The argument is moot simply because both styles are great, they just fulfill slightly different needs. When it comes to "Code as Documentation" its hard to argue against Protobuf, a format that is much easier for devs from other languages to buy into. I've never had a developer come to work on a Protobuf API who, after being shown the .proto files, could not understand how to read or extend the definitions.

I hope that developers will give the format a try because I think it's the next level up from normal web application design. It's the start of understanding that for larger applications, different tools should be considered to help alleviate the pains of a (potentially) larger system and the needs of moving data from one place to another on the fly.

Ok, that's a pretty intriguing statement. What different tools should we be looking at (or developing) to work on larger systems and larger data sets?

BJ: Hopefully I don't get myself into too much hot water with the answer to this question (or go off on a large tangent), but here we go. Keep in mind also that this long-winded answer comes with a grain of salt, because every system will be designed to meet different goals. Therefore, there is no "one true way" as some would tout.

That being said, if you are looking to build a system for growth, there are certain concepts and technologies that should at least be considered from the outset. Service-oriented Architecture (SOA) is a way of designing a system for growth, to me it's the most natural way to begin with the journey in mind. For those new to SOA, a short primer: It involves creating smaller independent applications that are easier to write and maintain because they focus on smaller feature sets, while when roped together you can gain the benefit of all the systems working as a whole and ready to scale.

In this type of system we never want to share data between service applications directly, such as connecting from Service A to Service B's database to get user data. We share data by creating APIs for each service application (with protobuf of course :)), then publish those APIs for our other services to consume. If one application needs user data, it doesn't connect to the user database, it connects to the internal User service's API to gather the data. Naturally protobuf fits extremely well here, but REST/JSON or SOAP or (insert other transport protocol here) can obviously be used also.

Other "large systems" or so-called "enterprise" technologies that fit well into an SOA system are background jobs (queues) and various types of messaging systems.

Queueing is essential for the speed and scalability of a system as it offloads non-relevant (yet important) processing to seperate threads or processes. A simple example of how a queue can give you an increase in speed and usability of a system is sending an email when a user is created. The user generally doesn't care (or know) that you are sending him an email when their account is created, but they do care that if its taking 10 seconds. So rather than tie up the user's process just to send an email, you would queue that "job" for later (even if it's processed milliseconds later) and let the process return the result of the user creation. Workers in other threads or processes will pick up the email job and send the email for you.

The main queueing system we use is Github's excellent Resque coupled with my own little resque-remote plugin. Resque-remote gives us the ability to queue a job for another service to consume.

Messaging is such an enormous topic that I'm not sure I'm the one you want to describe its ins and outs. The short of it is that in certain contexts we've found that it can make more sense to use push-based data transfer rather than pull-based. Take the user creation example: when a user is created in my User Service Application, the user service doesn't know about any other systems that may be interested that a user was created, and frankly it shouldn't care. The User Service should only be responsible to post a message (to a message service or bus) that an event occurred in the system, in this case a user was created. Once the event is messaged, user service creation can go about its merry way. Other parts of the system may be listening to the message (event) bus for user creation events and their associated data, and they will receive the data as a push. This specific messaging paradigm is usually referred to as PubSub (Publish/Subscribe). As I've already mentioned, there are many many more types of messaging patterns that can be followed.

These are just a few of the systems we've put in place to manage data transfer complexity in our SOA ecosystem. There's also another branch for data warehousing such as ETL data transfer systems like Pentaho or Jasper. The possibilities are... well, you get the idea.

The coolest part about all of this is that you can use Ruby for 100% of these so-called enterprise situations. We do. You don't have to use Java or .NET to solve "Big Boy" problems. When I first started with Ruby, I wasn't entirely sure of this, but I certainly am now.


So, you've read along this far. What do you think? How are you using Protocol Buffers? Why did you choose to go down this route?

Saturday, March 26, 2011

Ruby and Protocol Buffers, Take One and a Half

In a comment on my previous post on Protocol Buffers, Clayton O'Neill recommended trying out the java protobuf library with jruby. I'll get to that eventually, but his comment made me wonder how jruby and rubinius would do with this little test.
I fired up rvm and looped through my installed versions. Here are the results:
ruby 1.8.7 (2010-08-16 patchlevel 302) [i686-linux]
real 3m11.857s
user 3m11.024s
sys 0m0.124s
jruby 1.5.5 (ruby 1.8.7 patchlevel 249) (2010-11-10 4bd4200) (Java HotSpot(TM) Client VM 1.6.0_24) [i386-java]
real 2m54.035s
user 2m53.355s
sys 0m0.388s
rubinius 1.1.1 (1.8.7 release 2010-11-16 JI) [i686-pc-linux-gnu]
real 1m59.693s
user 2m5.292s
sys 0m0.148s
ruby 1.9.1p378 (2010-01-10 revision 26273) [i686-linux]
real 1m50.293s
user 1m49.811s
sys 0m0.092s
I certainly wouldn't choose a ruby implementation based on this alone, but it's good to see where things stand at the get go. As I keep going with this exploration, I'll try to keep posting timing results.

Update: Evan Phoenix (@evanphx) pointed out that I was using old versions of Rubinius and JRuby. Since I'm boxed into which Ruby implementation I use (1.8.7 on our boxes at work), I wasn't thinking about keeping things up to date in my RVM installation. I've updated JRuby and Rubinius and rerun the test. The results are as follows:
jruby 1.6.0 (ruby 1.8.7 patchlevel 330) (2011-03-15 f3b6154) (Java HotSpot(TM) Server VM 1.6.0_24) [linux-i386-java]
real 1m38.390s
user 1m42.914s
sys 0m0.508s
rubinius 1.2.4dev (1.8.7 536a6eb8 yyyy-mm-dd JI) [i686-pc-linux-gnu]
real 1m58.138s
user 2m1.492s
sys 0m0.144s

Thursday, March 24, 2011

Ruby and Protocol Buffers, Take One

At work, we're moving from XML to protocol buffers.  While we're mostly a Java shop, the operations/sysadmin team I'm on does a lot of Ruby. I was interested in how we might use the same technology for some of our stuff. After a bit of looking, I found two libraries that looked mature enough to investigate:

ruby-protobuf, by MATSUYAMA Kengo (@macks_jp), was straightforward to install and use.  It has a good online tutorial and the redme has all I needed to get started.
ruby-protocol-buffers, by Brian Palmer was also easy to install and use.  It seems a bit lacking in the online documentation, but does have some examples to follow.  (If Brian's name rings a bell, it might be because I interviewed him some time ago about winning a programming contest sponsored by Mozy's former incarnation.)
I started out with a very simple proto file:

package bench;

message Person {
  required string name = 1;
  required int32 id = 2;
  optional string email = 3;
}

I compiled this with rprotoc for ruby-protobuf and with ruby-protoc for ruby-protocol-buffers. This generated the following (which I edited lightly).  For ruby-protof:

### Generated by rprotoc. DO NOT EDIT!
### <proto file: bench.proto>
# package bench;
#
# message Person {
#   required string name = 1;
#   required int32 id = 2;
#   optional string email = 3;
# }
require 'protobuf/message/message'
require 'protobuf/message/enum'
require 'protobuf/message/service'
require 'protobuf/message/extend'

module Bench1
  class Person1 < ::Protobuf::Message
    defined_in __FILE__
    required :string, :name, 1
    required :int32, :id, 2
    optional :string, :email, 3
  end
end
for ruby-protocol-buffer
#!/usr/bin/env ruby
# Generated by the protocol buffer compiler. DO NOT EDIT!

require 'protocol_buffers'

# Reload support
Object.__send__(:remove_const, :Bench2) if defined?(Bench2)

module Bench2
  # forward declarations
  class Person2 < ::ProtocolBuffers::Message; end

  class Person2 < ::ProtocolBuffers::Message
    required :string, :name, 1
    required :int32, :id, 2
    optional :string, :email, 3

    gen_methods! # new fields ignored after this point
  end
end

Then I pulled out the statistical benchmarking I wrote about a while ago (Since no one else has taken the bait, maybe I should bundle up a gem for that.).  Instead of quoting the whole thing at you, here are the pertinent loops.  For ruby-protobuf:

msg = Bench1::Person1.new(:name => idx.to_s,
                          :id => idx,
                          :email => idx.to_s)
msg_str = msg.serialize_to_string
msg == msg.parse_from_string(msg_str)

for ruby-protocol-buffer:

msg = Bench2::Person2.new(:name => idx.to_s,
                          :id => idx,
                          :email => idx.to_s)
 msg == Bench2::Person2.parse(msg.to_s)
And here are the results:

$ rvm 1.9.2
$ ruby -v
ruby 1.9.2p0 (2010-08-18 revision 29036) [i686-linux]
$ time ruby ProtoBufBench
testing ruby-protobuff against ruby-protocol-buffer
The deviation in the deltas was 0.021731
The mean delta was 0.198301
max = 0.241761921640842 :: min = 0.154839326146633
ruby-protocol-buffer was better

real 1m50.672s
user 1m50.599s
sys 0m0.092s

$ rvm 1.8.7
$ ruby -v
ruby 1.8.7 (2010-08-16 patchlevel 302) [i686-linux]
$ time ruby ProtoBufBench
testing ruby-protobuff against ruby-protocol-buffer
The deviation in the deltas was 0.009414
The mean delta was -2.205984
max = -2.18715483485056 :: min = -2.22481263341116
There's no statistical difference

real 3m8.131s
user 3m7.996s
sys 0m0.056s

I didn't try compiling the c extension for ruby-protocol-buffers, and I haven't tried any more involved .proto files yet.  I'll work on those in the next couple of days and post results as I see them.

Wednesday, March 23, 2011

Review - Eloquent Ruby

The system management/administration team that I work on is starting to do more scripting and tool building.  That means bringing a bunch of people up to speed on Ruby.  We're using a combination of the Pickaxe Book and pair programming/mentoring to help bootstrap people.  So far it's been working pretty well.
Watching everyone else reading and learning made me want to get in on the action.  Fortunately, there's a new book out from Russ Olsen (@russolsen) — Eloquent Ruby.  I had the opportunity to interview Russ (look for it to show up soon) about his book, and Addison-Wesley was kind enough to send me a copy of it.
Eloquent Ruby is primarily aimed at people coming to Ruby from other languages.  It aims to explore and explain the idioms common to our community, and I think it does a great job of it.
It will serve current rubyists well too.  I learned several things from it as I read, and cemented other concepts as well.  Some of the book's explanations have crept their way into discussions about Ruby here at work.  It's good stuff.
Section One, on basics, is a rich source of little gems that will help you use Ruby's built in classes more effectively.  Section Two, on modules and classes, helps you build better classes of your own.  Section Three, which covers meta-programming, dives into an oft discussed but underused side of Ruby to push your Ruby-fu to the limit.  Section Four covers a variety of things that either don't fit into the other sections or build on concepts from them.
With this book, Russ has hit another one out of the park.  Go grab a copy for yourself.
If you're interested inDesign Patterns in Ruby  you can read my review here.  You can also read a previous interview with Russ.

Wednesday, September 08, 2010

Reading List Update 9/8/2010


The recent news that GDB now supports D makes The D Programming Language jump up a notch or two on my reading list.
I've finished 52 Loaves: One Man's Relentless Pursuit of Truth, Meaning, and a Perfect Crust, it was a fun read.  I really identified with his trip to the French monastery. It seemed like a great climax to his year, with the perfect denouement as he came home to bake his final loaves.
With a chance to get involved with a local restaurant group (not behind the counter though), the food books are still winning out in my what to read next decisions.

Tuesday, August 31, 2010

My Reading List on 8/31/2010


Thanks to Prentice Hall and Addison-Weseley giving me three new books, my reading list has bulked back up.  Here's what I'm working through at the moment:
What are you reading? Why?

Wednesday, August 25, 2010

Ruby|Web Interview with Pat Maddox


Ok, if I'm going to post about GoGaRuCo today, I should also spend some time on Ruby|Web, the latest regional conference from Mike Moore (@blowmage) and friends — truth in advertising, I'm a volunteer on the board for Ruby|Web, so I might be a bit biased.
Just so my biases don't show too much, I asked Pat Maddox(@patmaddox) to answer a few questions for me.  Of course, he's a speaker at Ruby|Web, so the spin is probably still there.
Regardless of our mind-set, this looks like it's going to be an awesome conference.  The only problem is that you need to register before Sep 3rd.  You don't have much time ... maybe you should go register first, then come back and read what Pat has to say.

Ruby|Web is a new name in the regional Ruby conference space.  What drew you to it?
Pat I love coding in Ruby, and I think the web is a great platform to develop for.  I'm eager to spend a few days rubbing elbows with like-minded people.  And between the over all theme, the fantastic organizers running the show, and Snowbird (!!), Ruby|Web jumped to the top of my list.  Plus I don't know anything about HTML 5 or CSS 3 and I need to get on that :)
Seaside is interesting technology, how did you discover it?
Pat Seaside is a fascinating and fun technology!!  I came across it a few years ago, not long after I got into Rails.  Over the years I had a couple of false starts with it...it's a bit opaque at first because the development environment is so different from anything I'm used to.  And it's only fairly recently with Pharo that it's become easier to get started, because it's such a clean environment geared towards development.  Also the documentation for both Pharo and Seaside are getting really good.  There are free books on each at pharobyexample.org and book.seaside.st/book.
Okay as for what's so interesting to me about Seaside... it's 50% the framework and 50% the Pharo environment. Seaside itself represents a step forward in web development similar to how Rails did.  Rails takes care of a lot of the plumbing for you - you don't have to parse query params, set up response headers, manage the session (unless you want to of course).  Seaside does all that of course but also manages application state for you.  So you don't have to worry about putting stuff into a database, then pulling it back out and operating on it.  I can't do it justice in a few sentences, but that's why I'll be showing lots of examples at the conference! :)  At any rate, that same feeling you get when you code Rails for the first time and see how much easier things are, you get that same feeling with Seaside.  It's not a replacement for Rails by any means - Rails definitely has a sweet spot, particularly when it comes to RESTful websites and interoperability with the unix ecosystem - but for the things that Seaside is strong at (which for me so far has been complex and/or configurable workflows), it runs circles around everything else.
The other thing I'm loving about Seaside development is Pharo, an open-source smalltalk environment.  Smalltalk is a great language, and Pharo has great tools that allow you to discover everything in the system.  Honestly it makes RubyMine or etags look plain silly.  The best bit is that nearly everything in Pharo is implemented in smalltalk, including all of the tools.  So if you want to see the mechanics of a refactoring tool, and even build your own, it's trivial to do so, because it's just smalltalk code.
Wow this answer got long.  I could go on all day about this stuff.  Gonna stop now.
What other smalltalk tools/ideas do you think Rubyists should be looking at?
Pat Let's see...I'd love to see Rubyists take the ideas from the Pharo IDE and build some really snazzy development environment for Ruby.  The next killer Ruby app, I think, is going to be a development environment that uses the runtime structure of objects to do all of its magic, rather than just statically analyzing source code.  Even just having portable refactoring tools would be awesome.
I'm also really excited about Maglev, whenever that becomes available for daily use.  It is incredibly liberating to write actual OO code, and so I think my style of coding Rails will change completely once Maglev enters the field.  fingers crossed
Which Ruby|Web presentations are you most looking forward to?

Pat In order of them being listed on the sessions page...
  • BJ Clark's (@robotdeathsquad) talk on HTML / CSS / Javascript.  He told me a few months back when he planned this talk that he thought, "if I were to school Pat on HTML / CSS / Javascript, what would I say?"  He and I have worked together for years and he gets frustrated with my lack of understanding of those things.  So basically this talk is geared specifically to people like me, hardcore backend developers with "div-itis" and who typically use inline javascript and CSS.  I'm looking forward to getting schooled.
  • Alistair Cockburn's (@TotherAlistair) samurai talk.  His talk summary means absolutely nothing to me (on purpose, I'm sure) but he's always a trip to watch speak, and I'm glad to see him get more exposure in the Ruby community.
  • Evan Light's (@elight) iOS talk - Evan is a diverse developer and entrepreneur.  Really excited to learn from his experiences.
  • Joe O'Brien's (@objo) communication talk.  For starters, Joe is one of my favorite people in the Ruby community.  Again, he's one of those folks that combines technical expertise with good business sense and a warm heart.  I think folks attending this conference are going to have more of the entrepreneurial spirit than most, so his talk will be particularly important and insightful for us.
  • Dirt Simple Datamining by Matthew Thorley (@padwasabimasala) - because really, who doesn't love datamining??
It's clearly shaping up to be a rocking conference!!!


GoGaRuCo 2010: mini-interview with Ilya Grigorik

Ilya Grigorik (@igrigorik) is another GoGaRuCo speaker who's kindly agreed to sit down and work through a short interview with me.  Hopefully this gives you taste of what you'll be missing if you're not going to the Bay Area's regional Ruby conference.

Machine Learning and Ruby don't leap to mind as a common pairing.  Why is machine learning important to Rubyists?
Ilya I don't think the topic of Machine Learning (ML) can or should be linked any specific language or runtime - it is much more general then that. Wikipedia provides a good starting point: "Machine learning is a scientific discipline that is concerned with the design and development of algorithms that allow computers to evolve behaviors based on empirical data". Artificial Intelligence is a close cousin to this definition, and you will find a lot of people using these two terms interchangeably, but I prefer the ML definition because, to me, defining and modelling the learning process is where the work happens, whereas "intelligence" is the ultimate outcome (plus, defining intelligence is a much harder concept to agree on).
With that in mind, I think you could make the argument that Rubyists apply ML to many everyday situations already: sorting algorithms, recommendations, and so on. It is also a truism that by the time any "AI" hits the mainstream, it is usually no longer interpreted as "AI". For example, computers answering telephone calls was the domain of pure science fiction only a few decades ago, whereas now we don't even stop to think about it. How about your ITunes "genius" playlist? Pandora, Last.fm? You see where I'm going. It's all around us.
Why is Ruby a good fit for machine learning applications?
Ilya Because Ruby commands such a presence in the web development world, I think it naturally finds itself in domains and applications that stand to gain a lot by leveraging the available data in some interesting and novel way. But once again, it's not really a question of language, as much as it is a question of modelling what you know, and applying that data to interesting questions. If Ruby, as a language, allows you to model your data in a faster or easier way, then so much the better.
On a purely practical, implementation side, Ruby has a number of great libraries and plugins that allow you to leverage many interesting algorithms: support vector machines, decision trees, bayes filters, neural nets, and so on. Will those tools scale to million row matrices? Perhaps not, but they will allow you to iterate through a number of solutions at a minimal cost, which in itself is a big win.
Where doesn't ruby fit well in the domain?
Ilya It is unlikely that you will be analyzing a multi-terabyte dataset with a Ruby ML algorithm. More likely, you'll work on a scale of a gigabyte (or a few), model and iterate your algorithm on a subset of data with Ruby, and then implement a lower level solution to scale up to larger datasets.
You're pretty well know for deep diving blogs about Ruby.  What's your day to day relationship with the language?
Ilya My day to day job is with PostRank, where being CTO/founder means I'm wearing many different hats throughout the day. Having said that, most of our systems are written in Ruby, and we have definitely pushed the limits of the language on many fronts.
My blog, is in many ways, a reflection of the technical challenges we're currently dealing with at PostRank, or technologies we're evaluating to improve our infrastructure. So, while I may not be working on implementing the next feature which is going into our analytics product, I am likely to be involved in the design and deployment of the infrastructure that has to deal with servicing all of the data requests required to make that feature possible (and when you're pushing as much data around as we do at PostRank, that's always a non-trivial challenge). The combination of having an awesome team, and a large and exciting problem to work on means there is never a shortlist of what to write about out on my blog!
Ilya
Other than your own talk, what are you most looking forward to at GoGaRuCo this year?
Ilya To be honest, every single talk on the agenda sounds fascinating to me - it's hard to pick any favorites. Having said that, I have been recently thinking and talking to a few people about the topic of "test driven learning", so I'm really looking forward to the "Test First Teaching" presentation by Sarah Allen and Alex Chaffee. I am really curious how this concept could be applied more broadly, outside of just learning a programming language. For example, could you structure a physics course in the same manner? Arguably some (great) teachers do this already, but I would love to extract and distill some general rules and patterns.

Thursday, August 19, 2010

GoGaRuCo 2010: mini-interview with Josh Susser


GoGaRuCo is just around the corner (Sep 17-18), and it looks like it's going to be a great conference again this year.  I wanted to touch base with Josh Susser (@joshsusser) again to see what was going to set this year apart.  He was kind enough to answer a few questions.  If you live in the Bay Area and haven't already decided to hit GoGaRuCo, what are you waiting for?

This is your second time around with GoGaRuCo.  What was your biggest lesson from 2009, and how did it impact this year's version?
JoshBiggest lesson?  Make sure you have comfortable chairs!  Well, chairs are important, but I think most about the technical program and how to put together the best one I can.  Last year we did a mix of invited talks and talks selected from submitted proposals.  And consistently, attendees rated the invited talks higher than ones selected from proposals.  Now, I can't say what that means overall and it's certainly not a reason for conferences to stop doing CFPs, but this year I wanted to try doing a curated program where all talks were invited.  That let us craft a program that has balance, diversity and high quality.  It can be more work than doing a CFP because you end up with some speakers who take a little more cat-herding than those that are self-motivated to submit a proposal, but it's worth it.  I know we might miss out on a great proposal from someone we wouldn't think to invite, but it also lets us have speakers who wouldn't think to submit a proposal, so I think it works out in the end.
Why the change from a Spring time slot to one in September?
JoshPlanning a conference takes a long time.  For GoGaRuCo we have to get started about six months in advance, and ramping up for an April conference meant we had to start around December, which is nearly impossible given that so many people are off on vacation or focused on wrapping up the old year or starting the new.  It made things terribly difficult to get off to a good start.  The hardest part of doing GoGaRuCo has been finding a venue that was large, affordable and good enough.  We just couldn't find a venue for April, so we moved later in the year to have more time.  It's also nice that September is the best weather of the year in San Francisco, so it will make the trip that much better for people visiting us.
The Wrap was a unique feature from GoGaRuCo last year.  How much impact has it had on the conference?  Will it be back this year?
JoshThe Wrap will definitely return this year.  The initial idea for The Golden Gate Ruby Wrap was to avoid having to print color program schedules.  We've seen them at a lot of confs, and they don't seem that useful and probably only exist as a place to print sponsor ads.  The Wrap saves on paper and cost, and gives us a more long-lived place for sponsor ads.  It also serves as a rich record of the conference, something like the proceedings published by traditional, academic conferences.  I don't know how much impact it's had, but the 2009 Wrap was downloaded by thousands of people, more even than looked at the videos of the talks.  I think that makes the historical impact of the conf greater, since there is an actual record of it.
If there's one reason people should come to GoGaRuCo this year, what is it?
JoshThis is San Francisco!  It's the highest concentration of hardcore Rubyists I know of.  Last year 96% of the conference attendees used Ruby as the primary language in their regular job.  That means being at the conf you are surrounded by people who love Ruby and know how to get things done with it.  And we have comfortable chairs.

Friday, July 16, 2010

Lone Star Ruby Conf Speaker Interview: Jesse Wolgamott


Today's a twofer for the Lone Star Ruby Conference.  My third interview (second today) is with Jesse Wolgamott (@jwo) who's presenting "Battle of NoSQL stars: Amazon's SDB vs Mongoid vs CouchDB vs RavenDB ".  Jesse shares some thoughts about NoSQL and the conference.



NoSQL looks like it's gaining momentum.  Why should Rubyists be interested in the topic?

Jesse Once you reach the point in transaction system where the database is the scalability cause of your scalability problems, there's no going back. You've taken the red pill. Table-based transaction databases are constrained by memory and there's a hard maximum until your app crawls to a halt. The dream of true replication and easy sharding is built in.

Also: migrations just suck, even in Rails.

There are a lot of NoSQL players, what made you zero in on these four?

Jesse CouchDB was my first experience with NoSQL -- the built-in map-reduce is so unique. MongoDB is a newer kid on the block, and is easier to get running in Rails, so that's a plus. I like to mention Amazon's SDB because it's frequently overlooked; You use Amazon's servers so in that respect it's the easiest to setup. RavenDB is new and shiny. Cassandra is also really cool, but not as ruby/rails friendly.

What's your current involvement in the NoSQL world?

Jesse I've used CouchDB, MongoDB, and SDB in the real world, but I'm just a lowly programmer.

What made you want to present at LSRC this year?

Jesse I got jealous looking at all the ruby conferences across the country, and heard about LSRC through the local Houston ruby group. Austin rules, Ruby rules, so win win win.

Other than your own talk, what are you most looking forward to at LSRC 2010?

Jesse The "Vim for the modern Rubyist" talk -- Vim is so hot right now! (but really, it should be cool and I love the AntiIDE thinking).

Lone Star Ruby Conf Speaker Interview: Nephi Johnson


Okay, time for a second interview with a Lone Star Ruby Conference speaker.  This time, Nephi Johnson (@d0c_s4vage) talks a bit about his presentation — "Less-Dumb Fuzzing and Ruby Metaprogramming".



Fuzzing isn't always well understood.  Can you describe fuzzing, and tell us what situations it's a good fit for?

Nephi Fuzzing is a term used to describe the process of feeding an application unexpected inputs in order to find flaws in the code.  During development, it's pretty much impossible to write code that will handle all possible inputs correctly.  Fuzzing helps to uncover some of the more subtle and unforeseeable flaws that haven't been found through code reviews and normal testing.  Fuzzing is typically very automated and usually involves feeding a program thousands or millions of sets of malformed input.  The program being fuzzed is then monitored for crashes, exceptions, and/or performance.

The last person to really talk about fuzzing in the Ruby Space was Zed Shaw.  That's kind of a tough act to follow.  Why is this an important topic for Rubyists, and why are you the right person to talk about it?

Nephi I think fuzzing is an important topic for every developer (or anyone who wants to find bugs in an application).  If you have the time to fuzz a product, you will almost certainly uncover flaws in it.  One bug fixed during development is one less bug that customers have to experience with a release product.  Not finding flaws in the code after extensive fuzzing is also a big confidence booster.  I think this is especially applicable to Rubyists because of the flexibility (and fun) that comes with using Ruby.  I think anyone could make their own fuzzer/data-generator with Ruby in a short amount of time.  Also, if someone wanted to use the library I've written, it just so happens that I've written it in Ruby [*sarcasm*].

Why am I the right person to talk about fuzzing?  Fuzzing is something that I spend most of my free time doing or working on.  I put a lot of thought into coming up with ways to more efficiently fuzz programs.  As a security researcher, I have different goals in fuzzing than developers.  I want to find the really interesting bugs, bugs that might allow one to run their own code or do something entirely unexpected with the program.  I think my perspective on fuzzing might provide different insights for those who use fuzzing outside of the security field.

What prompted you to speak at LSRC this year?

Nephi Someone had mentioned to me that it might be interesting to hear from somebody in the security field and suggested I submit a talk.  I liked the idea, chose to talk about a project I've been working on using Ruby, and here I am.

Other than your own talk, what are you most looking forward to?

Nephi I'm looking most forward to the talks "Vim for the modern Rubyist", "What every Ruby programmer should know about threads", and "Getting Started With C++ Extensions."  Why these talks?  I love using vim - it was the first text editor I used when I started using Linux and now it's all I use, threads give me trouble sometimes (ruby threads, that is), and I've been wanting to write my own ruby extensions for a while now.

Thursday, July 15, 2010

LSRC Speaker Interview with David Copeland


With the Lone Star Ruby Conference just over a month away, I thought it would be a good idea to talk to some of the presenters.  David Copeland (@davetron5000) is giving a talk about a topic that resonated with me, so I sent off an email to find out more about what he thought would make his presentation and the conference worthwhile.



I've never been a big 'web app' kind of guy, so I was excited to see "Why And How You Should Make Awesome Command Line Apps with Ruby" as a presentation.  Why do you think Ruby works well in this space?

Dave Having used PERL and bash in the past, Ruby is just a FAR more pleasant environment; it's just really easy to make a well-designed system, the code is clearer and easier to write, and there's a lot of great libraries that are easy to install and set up.   My talk doesn't go TOO heavily into this, but I think everything that makes Ruby great for web apps makes it great for command line apps.

As a Sys Admin, CLI stuff is my bread and butter.  Do you see Ruby as a good language for Sys Admins?  Why or why not?

Dave Nothing's going to be "as close to the system" as shell scripts, but Ruby has some great libraries that let you write cross-platform scripts, and that's a good thing.  Ruby also has a culture of terse-but-readable syntax, code-as-configuration and overall UNIXness that I think a sysadmin would find familiar and comforting.  Ruby really embodies the "motivated laziness" that is the hallmark of a good sysadmin (by which I mean automating painful tasks away into something simpler).  Further, there's some great management tools built with Ruby, like chef and capistrano.

What prompted you to present at the Lone Star Ruby Conference this year?

Dave I've spoken at a few conferences and user groups and really liked it, and I liked the idea of a Ruby-focused conference; Java-based conferences always feel a bit behind the bleeding edge to me, and the Ruby world is always pushing the boundaries.  I also thought my talk would be interesting to share, specifically for the reasons you note above; Ruby and Rails go hand in hand, but Ruby is an awesome language all on its own.  Plus, the last time I was in Austin, I was there for one night on a cross-country drive, so I didn't get to see much :)

Besides your own talk, what are you most looking forward to while you're there?

Dave There's a ton of really interesting looking talks scheduled; The ActiveModel/Active Relation talk looks good, as well as the NoSQL stuff and deployment talks.  And, I'm sure Tom Preston-Warner from github will give a good talk!