Live data from Hacker News

Big data: are we making a big mistake?

ft.com

11–20 of 90 posts

Re: Big data: are we making a big mistake?

#11
post #9

I find it amusing that the article talks about big mistakes in polling data, when the clear winner of the last two US elections is one Nate Silver, who aggregated polls to get predictions so close to the actual results, one wonders why people actually vote anymore. Now, just like with every other technological solution, we only learn about the limits of its use by overuse. There's plenty of people out there storing l…

But the article also talks about good polling data too. In effect, people have been making good and bad election predictions for decades.

Re: Big data: are we making a big mistake?

#12
Another conclusion to draw from this article (which I really enjoyed, by the way) is that Big Data has been turned into one of the most abstract buzzwords ever. You thought "cloud" was bad? "Big Data" is far worse in its specificity.

I can't count the number of times I'll be talking to some sales rep and they'll describe how they scan the data within whatever application they're demoing and "suggest" items using "big data techniques". In almost all cases they're talking about a few thousand or hundred thousand records, tops.

I've found that when non-hardcore techies talk about Big Data, what they really mean is "they have some data" vs before, when they had zero data.

From the article:

"Consultants urge the data-naive to wise up to the potential of big data. A recent report from the McKinsey Global Institute reckoned that the US healthcare system could save $300bn a year – $1,000 per American – through better integration and analysis of the data produced by everything from clinical trials to health insurance transactions to smart running shoes.

What these consultants mean is that by having just some data compared to the silo'd data that is the norm in US healthcare, they could save a lot, and they're right. My previous company had a large data set (20+ million patients) and we'd find millions of dollars of savings opportunities for every hospital we implemented in, but that's because we had the data, not because we were running some kind of non-causual correlation analysis like the article references. It was just because we could actually run queries on a data set.

-----

Off Topic - how annoying is it that when you copy & paste from the FT, they preface your copy with the following text?

High quality global journalism requires investment. Please share this article with others using the link below, do not cut & paste the article. See our Ts&Cs and Copyright Policy for more detail. Email ftsales.support@ft.com to buy additional rights. http://www.ft.com/cms/s/2/21a6e7d8-b479-11e3-a09a-00144feabd...

Re: Big data: are we making a big mistake?

#13
This article reminds me of the argument [0] between Noam Chomsky [1] and Peter Norvig [2]. TL;DR (paraphrased with hyperbole) Chomsky claims the statistical AI of Norvig is a fancy sideshow that doesn't understand _why_ it is doing a thing. It just throws gigabytes of data at an ensemble and comes out with an answer.

[0] - http://www.theatlantic.com/technology/archive/2012/11/noam-c...

[1] - http://en.wikipedia.org/wiki/Noam_Chomsky

[2] - http://en.wikipedia.org/wiki/Peter_Norvig

----

Norvigs rebuttal, http://norvig.com/chomsky.html

Re: Big data: are we making a big mistake?

#14

Another conclusion to draw from this article (which I really enjoyed, by the way) is that Big Data has been turned into one of the most abstract buzzwords ever. You thought "cloud" was bad? "Big Data" is far worse in its specificity. I can't count the number of times I'll be talking to some sales rep and they'll describe how they scan the data within whatever application they're demoing and "suggest" items using "big…

Out of curiosity, when does it effectively become "big data"?

I ask not to be snarky, but it might be the case that it's "big data" to someone else, but not necessarily to you. I figured it was a relative term for your industry/business, but the hacker crowd definitely seems to peg that amount in the millions of data points before calling it big data at all.

Seems fair, but I'd rather clarify.

Re: Big data: are we making a big mistake?

#15
post #14

Another conclusion to draw from this article (which I really enjoyed, by the way) is that Big Data has been turned into one of the most abstract buzzwords ever. You thought "cloud" was bad? "Big Data" is far worse in its specificity. I can't count the number of times I'll be talking to some sales rep and they'll describe how they scan the data within whatever application they're demoing and "suggest" items using "big…

Out of curiosity, when does it effectively become "big data"? I ask not to be snarky, but it might be the case that it's "big data" to someone else, but not necessarily to you. I figured it was a relative term for your industry/business, but the hacker crowd definitely seems to peg that amount in the millions of data points before calling it big data at all. Seems fair, but I'd rather clarify.

Big Data used to mean petabytes, ie above the limits of performant scale up.

Re: Big data: are we making a big mistake?

#16
post #6

> a provocative essay published in Wired in 2008, “with enough data, the numbers speak for themselves” I think that's indicative of Wired breathless enthusiasm for technology that turned my off buying the print version many years ago. Scrape away some of the hyperbole and it is true that data driven management has made many companies more competitive and, if I dare mention the hobgoblin, efficient. Hunches and ideas…

I have some Wired issues from mid 90s in the bathroom and the tone is the same.

It seems pretty much everything they write about is supposed to change the world in a major paradigm shift.

Re: Big data: are we making a big mistake?

#17
post #8

if i work for facebook and i want to figure out something about my users, isn't it safe to say N = All since the data im accessing is all user data from fb? it's easy to go wrong with big data, and although the article glossed over some fairly important things (assuming the people who work on these datasets are much dumber than they are in reality), they're right on about idea that the scope and scale of what big dat…

Whilst, in the example you provide, it might be the case that "N = all", the cautionary tale offered in the article is that you always need to make sure you are asking the right question, and it is pretty easy to confuse yourself.

So you said "if i work for facebook and i want to figure out something about my users", and for whatever you were doing, looking at your existing user base might be the right thing to do. Perhaps, though, you actually want to know something about all your potential users, not just the users you happen to have right now. Whether or not your current user base offers a good model for your potential user base would then be a pretty important question, and one that almost certainly isn't answered by "big data".

I think that, as with most of statistics, the key point is "think about your problem", and that focusing on a set of solutions rather than the problems themselves can get in the way of that.

Re: Big data: are we making a big mistake?

#18
Nonsense. Google Flu was not "Big" data, they had only a few years worth of data at best. Additionally, when combined with current CDC data, it's predictions were better than models based on CDC data alone. And in all likelihood they can improve it with better methods.

Re: Big data: are we making a big mistake?

#19
A few other comments have raised this point, but Big Data is basically the new Web 2.0. Aside from being a buzzword, as a term it's so nebulous that half of the articles about it don't really define what it is. When does "data" become "big data"?

Re: Big data: are we making a big mistake?

#20
post #9

I find it amusing that the article talks about big mistakes in polling data, when the clear winner of the last two US elections is one Nate Silver, who aggregated polls to get predictions so close to the actual results, one wonders why people actually vote anymore. Now, just like with every other technological solution, we only learn about the limits of its use by overuse. There's plenty of people out there storing l…

Nate's approach is based on evaluating the quality of the various polls - which is the thrust of the FT article. In fact he actively weighted each of the polls & corrected for known biases.
Post reply on HN