Live data from Hacker News

Hadoop Reaches 1.0

hadoop.apache.org

1–10 of 31 posts

Re: Hadoop Reaches 1.0

#5
Hadoop reaches 1.0 and my understanding of how to use it is still in development.

Does anyone have a high level resource of how MapReduce works for mediocre programmers like myself that are late to the game? I know she's not ready to have my babies, but surely I could get to know her a little, maybe just be friends? I grabbed a Hadoop pre-made virtual machine the other month and was surely so far over my head that I had to run away to regroup.

In general I have some very unoptimized problems that MapReduce probably isn't the right shoe for, but I'd love to explain to my boss why it's the wrong shoe. And learning about it might be a great start down that path.

Re: Hadoop Reaches 1.0

#7

It was already prod ready in my opinion. I think this release is more of a "polish" thing since some people are timid to run "0.20" code in prod.

Agreed.

Working with hadoop a few years ago was a pain in the ass, what really made it ready (at least for me) was the packaging done by Cloudera.

Re: Hadoop Reaches 1.0

#8

Hadoop reaches 1.0 and my understanding of how to use it is still in development. Does anyone have a high level resource of how MapReduce works for mediocre programmers like myself that are late to the game? I know she's not ready to have my babies, but surely I could get to know her a little, maybe just be friends? I grabbed a Hadoop pre-made virtual machine the other month and was surely so far over my head that I…

A good introduction to MapReduce is probably CouchDB, where you use it for database views instead of SQL-style queries. The basic concepts are:

- The "Map" phase takes a key/value pair of input and produces as many other key/value pairs of output as it wants. This can be zero, it can be one, or it can be over 9000. Each Map over a piece of input data operates in isolation.

- The "Reduce" phase takes a bunch of values with the same (or similar, depending on how it's invoked) keys and reduces them down into one value.

A good example is, say you have a bunch of documents like this:

    {"type": "post",
     "text": "...",
     "tags": ["couchdb", "databases", "js"]}
And you want to find out all the tags, and how many posts have a given tag. First, you have a map phase:

    function (doc)
      if (doc.type === "post") {
        doc.tags.forEach(function (tag) {
          emit(tag, 1);
        });
      }
    }
In this case, it filters out all the documents that aren't posts. It then emits a `(tag, 1)` pair for each tag on the post. You may end up with a pair set that looks like:

    ("c", 1)
    ("couchdb", 1)
    ("databases", 1)
    ("databases", 1)
    ("databases", 1)
    ("js", 1)
    ("js", 1)
    ("mongodb", 1)
    ("redis", 1)
Then, your reduce phase may look like:

    function (keys, values, rereduce) {
      return sum(values);
    }
Though the kinds of results you get out of it depend on how you invoke it. If you just reduce the whole dataset, for example, you get:

    (null, 9)
Because that's the sum of the values from all the pairs. On the other hand, running it in group mode will reduce each key separately, so you get this:

    ("c", 1)
    ("couchdb", 1)
    ("databases", 3)
    ("js", 2)
    ("mongodb", 1)
    ("redis", 1)
Since the sum of all the pairs with "databases" was 3, the value for the pair keyed as "databases" was 3. You're not limited to summing - any kind of operation that aggregates multiple values and can be grouped by key will work as well.

Like you said, there are problems that this doesn't work for. But for the problems it does work for, it's very computationally efficient and fun.

Re: Hadoop Reaches 1.0

#9
Awesome! Now if we can just get HBase to update it's prereqs and bump it's version, I can have some symmetry in my life!

On a more serious note - is anyone using HDFS for something like the WebHDFS stuff was designed? We're currently looking at HDFS right now for an Event Store mechanism, but it appears to me to be pretty large file / stream oriented, and I'm wondering how it will stack up if we want to do something that involves files much smaller than say, 64MB.

Re: Hadoop Reaches 1.0

#10

It was already prod ready in my opinion. I think this release is more of a "polish" thing since some people are timid to run "0.20" code in prod.

Agreed. Working with hadoop a few years ago was a pain in the ass, what really made it ready (at least for me) was the packaging done by Cloudera.

I've run clusters with and without Cloudera and I'd never go back to without it. It just works which when setting up a new cluster is often something you can't say.
Post reply on HN