Live data from Hacker News

Inside Google Brain

wired.com

41–50 of 77 posts

Re: Inside Google Brain

#41

"They’ve also found that the models tend to become more accurate the more data they consume. That may be the next big goal for Google: building AI models that are based on billions of data points, not just millions. " I'm not versed in machine learning, but it looks to me that any model whose output quality is dependent on the quantity of data it ingests is deeply flawed. There's no doubt a bigger number of samples w…

It may seem counterintuitive, but that's a relatively common result in machine learning. Often the issue isn't so much the quantity of data as the representative nature of your training data used to create the model. All other things being equal, a larger set of training data is likely to be more consistent with the true data. Caveats abound, but that's the general idea.

One way to think about it is to look at problems with human perception like forced perspective. It's relatively easy to create a situation where the only available information results in mental models that describe the size of an object incorrectly. Given a different point of view (i.e. more information) the faults in the model become obvious.

Re: Inside Google Brain

#42

"They’ve also found that the models tend to become more accurate the more data they consume. That may be the next big goal for Google: building AI models that are based on billions of data points, not just millions. " I'm not versed in machine learning, but it looks to me that any model whose output quality is dependent on the quantity of data it ingests is deeply flawed. There's no doubt a bigger number of samples w…

You're sort of describing the problem of "over fitting", which is now very well understood in machine learning circles. That's when you get a model that describes it's training data very well, but doesn't generalise well.

The thing about using lots of data is that prior to publication of "The Unreasonable Effectiveness of Data" in 2009, most people did think that good algorithms were the most important thing. What that research showed was that a bad algorithm given more data will eventually outperform "better" algorithms, at least when those algorithms are initially judged based on their performance on smaller datasets.

So what happened with neural nets was that after some initial excitement about how they were more like the human brain, etc., it was found that "stupider" algorithms actually performed better and ANNs were written off for a while. It turns out that the reason NNs were performing badly was that they weren't being fed enough data.

Nowadays it's pretty easy to saturate a feed-forward neural network with data to the point where it's performance will never get much better. Deep learning techniques allow you to train bigger and more complex models with more data, but these more complex neural nets won't perform very well unless you feed them tons of data.

So in reference to your point about the brain, the thing about brains is that they actually learn based on massive amounts of data too. Think about how much data you have from continually streaming video ~16 hours/day, plus sound, plus touch, proprioception, and other inputs, over the course of many years.

Deep learning tries to emulate this to some degree with "pre training" which is where you feed lots of data into a deep network and have it learn "something" (it learns by itself at this stage). Then you start teaching it more complicated, high-level concepts. This pre-training allows it to do things like recognise common patterns in images, which the later training allows it to then associate with semantic ideas like "this is an apple", "this is a person", etc.

TL;DR: What seems to work best is fairly "dumb" algorithms, scaled up to be able to handle vast amounts of information and fed a ton of data to learn from. This is also how the human brain works.

Re: Inside Google Brain

#43
post #40
post #37

"Google isn't really a search company-- it's a machine learning company." No, it's an advertising company. That's who pays the bills. At the end of the day all this cool tech is to better understand and model human beings in order to better push ads. Sometimes it depresses me that so many of the world's most brilliant minds are working on that, but a generation or two ago they'd all be building doomsday bombs. I gues…

If there was anything else they could do that was profitable then they would. Do you have any suggestions? The only reason Apple has as much money as it does is because it greatly overcharges customers, doesn't participate in research that benefits society, and takes advantage of it's customers psychological need to have the latest model (even if the improvements are minimal).

I wasn't even really dissing Google, just pointing out the reality of the world we live in. The fact is that very very few people really care about visionary things. There has to be a way to make it pay, and unfortunately surveillance-based marketing is the only big meal ticket in town for services like Google.

Re: Inside Google Brain

#45
post #43
post #40

Earlier quoted context omitted.

If there was anything else they could do that was profitable then they would. Do you have any suggestions? The only reason Apple has as much money as it does is because it greatly overcharges customers, doesn't participate in research that benefits society, and takes advantage of it's customers psychological need to have the latest model (even if the improvements are minimal).

I wasn't even really dissing Google, just pointing out the reality of the world we live in. The fact is that very very few people really care about visionary things. There has to be a way to make it pay, and unfortunately surveillance-based marketing is the only big meal ticket in town for services like Google.

It is true about the ads being their primary source of revenue, I've thought long and hard about the concept but eventually accepted the ads as being necessary to fund the other projects they work on.

Their research isn't solely focused on gathering data for ads. If these technologies can be applied to ads then it helps justify large costs because it'll increase revenue, but that is not the sole driving motivator for their moonshots and research.

Hopefully the percentage of people who are passionate about futuristic concepts and visionary ideas will rise in the near future. The faster we get to a post-scarcity society the greater chance we, as an intelligent species, will have to survive long-term.

Just my 2 cents. I love them as a company and have a huge amount of respect for what they've done and for sticking to their founding principals of transparency and "do no evil" (for the most part, within reason for a company of their size).

I know it's pretty subjective and irreverent to bring Apple into any Google debate, the amount of people who put them on a pedestal drives me crazy. Anything I can do to redirect money from going to Apple is a positive thing in my opinion.

Re: Inside Google Brain

#47

Earlier quoted context omitted.

Your first statement here is not true. Humans are excellent at learning from very few or even 1 example. Show a toddler a single image of an elephant and the toddler will generalize perfectly on new examples; show a machine a few thousand images of elephants and it might generalize decently if your machine is really clever. There are very few tasks where machine systems achieve anything resembling human level perform…

That toddler has already processed lots of visual image data, examples of objects, nonliving and living, animals, mammals, etc. Don't you think that constitutes a large, important dataset for the problem of elephant recognition?

Yeah i agree. When the child is shown the labeled example of an elephant and infers the traits that make an elephant an elephant, her previous visual experiences provide background knowledge that restricts the space of hypotheses she considers. After all, there's an infinite set of logically consistent hypotheses.

Nevertheless, if you provide your machine system with the video of all the child's visual input, it still won't generalize well from single examples, the way children do effortlessly.

Re: Inside Google Brain

#49

Earlier quoted context omitted.

That toddler has already processed lots of visual image data, examples of objects, nonliving and living, animals, mammals, etc. Don't you think that constitutes a large, important dataset for the problem of elephant recognition?

Yeah i agree. When the child is shown the labeled example of an elephant and infers the traits that make an elephant an elephant, her previous visual experiences provide background knowledge that restricts the space of hypotheses she considers. After all, there's an infinite set of logically consistent hypotheses. Nevertheless, if you provide your machine system with the video of all the child's visual input, it stil…

This reminds me of the saying: It took me ten years to become an overnight success.

Humans generalize well from few examples because, well, they've already processed billions of examples. A toddler may have never seen an elephant before, but it may have seen cars, trucks, birds, dogs, people, trees, skies, buildings etc, giving it concepts for bigness, smallness, aliveness, humanness and much else. With all these concepts in place, then yes it becomes easy to see what makes an elephant distinct from a dog or a person. And it would be too for an artificial neural network.

An interesting fact is that newborns have very few concepts to begin with. It takes some months for them for instance to learn to differentiate between alive and dead things (the family cat vs a teddy bear for instance).

Post reply on HN