Live data from Hacker News

Big data: are we making a big mistake?

ft.com

31–40 of 90 posts

Re: Big data: are we making a big mistake?

#31
post #30
post #14

Earlier quoted context omitted.

Out of curiosity, when does it effectively become "big data"? I ask not to be snarky, but it might be the case that it's "big data" to someone else, but not necessarily to you. I figured it was a relative term for your industry/business, but the hacker crowd definitely seems to peg that amount in the millions of data points before calling it big data at all. Seems fair, but I'd rather clarify.

When you are constrained to O(n) methods, you have big data.

Some people are constrained to O(WTF) methods and have no idea about O. So everything is Big Data.

Re: Big data: are we making a big mistake?

#32
The “with enough data, the numbers speak for themselves” statement has several meanings.

In one sense, if you can observe real phenomena, you don't have to guess at what is happening. For businesses that collect troves of it, they may need statistics 'less' because the sample size may approach the population size.

But calculating basic (mean, standard deviation, etc.) statistics is hardly the most interesting part. Inferential statistics is often more useful: how does one variable affect another?

As the article points out, the "... the numbers speak for themselves” statement may also be interpreted as "traditional statistical methods (which you might call theory-driven) are less important as you get more data". I don't want to wade in the theory-driven vs. exploratory argument, because I think they both have their places. Both are important, and anyone who says that only one is important is half blind.

Here is my main point: data -- in the senses that many people care about; e.g. prediction, intuition, or causation -- does not speak for itself. The difficult task of thinking and reasoning about data is, by definition, driven by both the data and the reasoning. So I'm a big proponent of (1) making your model clear and (2) sharing your model along with your interpretations. (This is analogous to sharing your logic when you make a conclusion; hardly a controversial claim.)

Re: Big data: are we making a big mistake?

#33
post #28
post #14

Earlier quoted context omitted.

Out of curiosity, when does it effectively become "big data"? I ask not to be snarky, but it might be the case that it's "big data" to someone else, but not necessarily to you. I figured it was a relative term for your industry/business, but the hacker crowd definitely seems to peg that amount in the millions of data points before calling it big data at all. Seems fair, but I'd rather clarify.

I usually follow DevOps Borat's definition [1]: "Big Data is any thing which is crash Excel." Many a true word spoken in jest. [1] https://twitter.com/DEVOPS_BORAT/status/288698056470315008

"Small Data is when is fit in RAM. Big Data is when is crash because is not fit in RAM."

https://twitter.com/DEVOPS_BORAT/status/299176203691098112

Re: Big data: are we making a big mistake?

#34
post #28
post #14

Earlier quoted context omitted.

Out of curiosity, when does it effectively become "big data"? I ask not to be snarky, but it might be the case that it's "big data" to someone else, but not necessarily to you. I figured it was a relative term for your industry/business, but the hacker crowd definitely seems to peg that amount in the millions of data points before calling it big data at all. Seems fair, but I'd rather clarify.

I usually follow DevOps Borat's definition [1]: "Big Data is any thing which is crash Excel." Many a true word spoken in jest. [1] https://twitter.com/DEVOPS_BORAT/status/288698056470315008

This is very inaccurate/misleading IMHO. Big Data is something which does not fit in a regular machine for a given operation. You can sort billions of records on an iPhone, for example. You can grep a string within a terabyte-file data on a single personal computer, and I am not convinced you'd go faster with a distributed system (reading the file on cold storage will be the limiting factor). People claiming to do "big data" in these situations do not generally understand the underlying concepts.

Re: Big data: are we making a big mistake?

#35

The misconceptions about big data are similar to those surrounding the word science. Many people associate "science" with things: cells, microscopes, the inner workings of the body. But science isn't a set of things; it's a process, a method of thinking, that can be applied to any facet of life. Big data is similar, in my opinion. It's not so much about the stuff — the size or diversity of a company's datasets. It ha…

It's the misconception that measurable observations equal the real distribution of the underlying events. Even professional data people often get that wrong, and it's not strictly limited to big data.

One of the most obvious examples was this one: A data set of all known meteorite landings[1] turns into "Every meteorite fall on earth mapped" [2] with looks like a world population maps sprinkled with some deserts known for their meteorite hunter tourism. The actual distribution can be theoretically described as a curve falling towards the poles.[3]

While this example is pretty obvious, one could expect similar observation biases in other data sources. A danger lies where data analyst do not bother to investigate what their data actually represents and then go on to present their conclusions like it would be some kind of universal truth.

[1]http://visualizing.org/datasets/meteorite-landings

[2]http://www.theguardian.com/news/datablog/interactive/2013/fe...

[3]http://articles.adsabs.harvard.edu//full/1964Metic...2..271H...

previous discussion of this: https://news.ycombinator.com/item?id=5240782

Re: Big data: are we making a big mistake?

#36
post #13

This article reminds me of the argument [0] between Noam Chomsky [1] and Peter Norvig [2]. TL;DR (paraphrased with hyperbole) Chomsky claims the statistical AI of Norvig is a fancy sideshow that doesn't understand _why_ it is doing a thing. It just throws gigabytes of data at an ensemble and comes out with an answer. [0] - http://www.theatlantic.com/technology/archive/2012/11/noam-c... [1] - http://en.wikipedia.org/w…

I think they are both wrong. You need better models than just throwing lots of data at something simple which Norvig likes. But they are still statistical models at some level.

Re: Big data: are we making a big mistake?

#38

Another conclusion to draw from this article (which I really enjoyed, by the way) is that Big Data has been turned into one of the most abstract buzzwords ever. You thought "cloud" was bad? "Big Data" is far worse in its specificity. I can't count the number of times I'll be talking to some sales rep and they'll describe how they scan the data within whatever application they're demoing and "suggest" items using "big…

I enjoyed your post, nemesisj. Within your field of longitudinal patient data, if I am correct in what you have written, your large datasets simply have a new name, paranthetically Big Data, and that you could get what you need to save money without the newfangled algorithms. Within academic bioscience, I think there is great consensus on what Big Data is - I have not seen much argument at all; however, it is still very hard to define within that field. The best I can do, over this cup of coffee, is to state that there is a clear distinction between the study of a gene, up to a few pathways vs. computational analysis of multiple OMICS (genomics, metabolomics and proteomics) datasets. I know that definition is terribly lacking and I am fighting the urge to delete it for the sake of getting the post completed. Anyway, Big Data is clearly changing the academic biosciences through the funding trend. That is, grants with a computational focus, or sub focus, certainly seem to be doing comparably well. I mention this because todays academic funding trends influence the direction of tomorrows startups as those being trained are disproportionally within the better funded labs, and draw upon their previous experience when forming companies. So, I personally believe this Big Data thing, however it is best defined over all, or within a given field, is in some way something new, and will continue to shape the startup sphere for years to come, especially in the areas of genomics, metabolomics and proteomics. This is my first post: ) Thanks!

Re: Big data: are we making a big mistake?

#39

Another conclusion to draw from this article (which I really enjoyed, by the way) is that Big Data has been turned into one of the most abstract buzzwords ever. You thought "cloud" was bad? "Big Data" is far worse in its specificity. I can't count the number of times I'll be talking to some sales rep and they'll describe how they scan the data within whatever application they're demoing and "suggest" items using "big…

Noam Chomsky had the best response to "big data": it's basically a nonsense concept (which I agree with) because "thinking is hard."

Aiming to support the parent's claim (ElDiablo66, is this what you're referring to?), here's a quote from an article [1] about the Chomsky vs. Norvig debate a couple years ago:

Chomsky critiqued the field of AI for adopting an approach reminiscent of behaviorism, except in more modern, computationally sophisticated form. Chomsky argued that the field's heavy use of statistical techniques to pick regularities in masses of data is unlikely to yield the explanatory insight that science ought to offer. For Chomsky, the "new AI" -- focused on using statistical learning techniques to better mine and predict data -- is unlikely to yield general principles about the nature of intelligent beings or about cognition.

[1] http://www.theatlantic.com/technology/archive/2012/11/noam-c...

[2] HN thread on [1]: https://news.ycombinator.com/item?id=4729068

Re: Big data: are we making a big mistake?

#40
post #13

This article reminds me of the argument [0] between Noam Chomsky [1] and Peter Norvig [2]. TL;DR (paraphrased with hyperbole) Chomsky claims the statistical AI of Norvig is a fancy sideshow that doesn't understand _why_ it is doing a thing. It just throws gigabytes of data at an ensemble and comes out with an answer. [0] - http://www.theatlantic.com/technology/archive/2012/11/noam-c... [1] - http://en.wikipedia.org/w…

Norvig's rebuttal is an excellent read.
Post reply on HN