What is the author of this article trying to say here? > it appears that for both Facebook and Yahoo, those same clusters are unnecessary for many of the tasks which they’re handed. In the case of Facebook, most of the jobs engineers ask their clusters to perform are in the “megabyte to gigabyte” range (pdf), which means they could easily be handled on a single computer—even a laptop. That facebook or yahoo could be…
Most data isn’t “big,” and businesses are wasting money pretending it is
51–60 of 160 posts
Re: Most data isn’t “big,” and businesses are wasting money pretending it is
#52This article is the equivalent of "horse drawn carriages are perfectly adequate for most journeys, and much more pleasant and commodious to boot." Good luck with that, buddy. You're not going to know what correlations are important and which are not until you study the data. Telling people to just collect the "important data" is like telling someone who has lost his keys just to go back to where he left them. It's al…
Re: Most data isn’t “big,” and businesses are wasting money pretending it is
#53As some one who is currently dealing with these sort of things I can tell this article hits the nail on its head. Most, heck something like 99.99% of all so-called big data I've dealt is something I wouldn't even classify as small data. I've seen data feeds in KB's sent over to be handled in as big data. It happens all the time. A simple data problem sufficient enough to be easily solved on something like a small db…
Re: Most data isn’t “big,” and businesses are wasting money pretending it is
#54If I ever want to get rich, I'll set up shop convincing small businesses they need to do things the way Google does, if only they want to remain competitive. Oracle has used exactly this business model to great success, and obscene profit, for over 30 years.
Edit: I noticed you meant small businesses. However, Oracle does this mainly for the large companies that don't excel at technology.
Re: Most data isn’t “big,” and businesses are wasting money pretending it is
#55Earlier quoted context omitted.
dewitt said "Oracle has used exactly this business model to great success, and obscene profit, for over 30 years". You claim that business models are abstract. Perhaps, except in the case where the words "exact" are used, and a specific company is named. I, like SeoxyS find the paradox of Oracle being accused of helping businesses to be "me-too" copies of Google - a company half of Oracle's age - amusing.
The model in question is "convince small and medium businesses that they need to buy my software in order to do things the same way as large companies X and Y and have a hope of remaining competitive". For Oracle X and Y were banks, retail and logistics companies, for the new generation of "big data" vendors it is Google and Facebook.
Re: Most data isn’t “big,” and businesses are wasting money pretending it is
#56Earlier quoted context omitted.
I know that's not what you mean, but I find it quite amusing that you describe Oracle (a 30 year old company)'s business as convincing people they need to do things the same way Google (a 15 year old company) does it.
To be pedantic, he said Oracle follows the same business model , not that Oracle follows Google .
Re: Most data isn’t “big,” and businesses are wasting money pretending it is
#57Nonetheless, large jobs are important too. Over 80% of the IO and over 90% of cluster cycles are consumed by less than 10% of the largest jobs (7% of the largest jobs in the Facebook cluster). These large jobs, in the clusters we considered, are typically revenue-generating critical production jobs feeding front-end applications.
So MR job characteristics might follow a power law distribution, and @mims is focusing on one end of the tail. Sure, that's cool!
But then @mims also selectively quotes the TC article, which ends with an excellent point that contradicts his thesis:
The big data fallacy may sound disappointing, but it is actually a strong argument for why we need even bigger data. Because the amount of valuable insights we can derive from big data is so very tiny, we need to collect even more data and use more powerful analytics to increase our chance of finding them.
I think @mims over-pursues the stupid Forbes/BI straw man here. As one would expect with data, the story is complicated. Mom and pop stores don't need to worry about Cloudera's latest offering, but companies working on the cutting edge of analysis still absolutely need tools like Hadoop, Impala, and Redshift.
Re: Most data isn’t “big,” and businesses are wasting money pretending it is
#58Re: Most data isn’t “big,” and businesses are wasting money pretending it is
#59This article is the equivalent of "horse drawn carriages are perfectly adequate for most journeys, and much more pleasant and commodious to boot." Good luck with that, buddy. You're not going to know what correlations are important and which are not until you study the data. Telling people to just collect the "important data" is like telling someone who has lost his keys just to go back to where he left them. It's al…
He doesn't tell anyone to collect the "important data," and he doesn't insist FB or Yahoo are not web scale.
His concluding paragraph is relatively weak, but the main thesis -- most businesses can ignore the Forbes/BI crap and analyze their data sufficiently using normal tools -- is true and sound.
Re: Most data isn’t “big,” and businesses are wasting money pretending it is
#60For me, "big" data is increasing the linkage between your data. It's not simply more data, but much richer, less formal data relationships. It's taking your sales data and linking it to your website clicks, linking that to the weather (or whatever). Or you take something traditionally static and add a temporal dimension. This kind of deep linking you can't measure with straight megabytes. A few gig doesn't seem that…