Live data from Hacker News

Marc Andreessen on the atomization of AI

techcrunch.com

21–30 of 47 posts

Re: Marc Andreessen on the atomization of AI

#21

The common saying is "more data usually beats a better algorithm". However, this is only for short term progress. Long term progress in the field of AI clearly requires better algorithms, and doing more with less data is exactly the kind of problem that a startup in the field could solve with a clever idea. That said, these startups will have to be extremely research heavy — how does one draw the most talented resear…

>> doing more with less data is exactly the kind of problem that a startup in the field could solve with a clever idea. It will take much more than a "clever idea" to overcome the reliance on data of the current AI state of the art- that is to say, machine learning. If you think about it, machine learning algorithms are essentially clever search procedures for some optimum in a heap of data (that's optimisation algor…

> There's probably some sort of way to make a system that can learn from very few examples and at a very short time, like we do ourselves.

Maybe. I wonder if we can pick up identifying platypus with only a couple of examples, because we spent a long time identifying bird or cat as a toddler. We have a complicated model of the world, and we can take short cuts, because we just identify the differences compared to stuff we already know.

You may be right, no one knows for sure. I think it's one of those P = NP kinda problems. Lots of clever people have thought about it for years. Maybe there's a trick. It seems unlikely there's a trick though. Absence of evidence starts to become evidence of absence.

Re: Marc Andreessen on the atomization of AI

#22

The common saying is "more data usually beats a better algorithm". However, this is only for short term progress. Long term progress in the field of AI clearly requires better algorithms, and doing more with less data is exactly the kind of problem that a startup in the field could solve with a clever idea. That said, these startups will have to be extremely research heavy — how does one draw the most talented resear…

> how does one draw the most talented researchers to work for a small group of people with minimal pay

Universities seem to be doing alright at this.

Probably the trick is to offer an exciting problem and lots of freedom.

Re: Marc Andreessen on the atomization of AI

#23

The common saying is "more data usually beats a better algorithm". However, this is only for short term progress. Long term progress in the field of AI clearly requires better algorithms, and doing more with less data is exactly the kind of problem that a startup in the field could solve with a clever idea. That said, these startups will have to be extremely research heavy — how does one draw the most talented resear…

I have to fundamentally disagree with Marc here.

Someone who actually DOES AI is Fei Fei Li at stanford. She's focusing on creating datasets like imagenet and more recently the visual genome: https://www.technologyreview.com/s/545906/next-big-test-for-...

A lot of AI startups' best assets are their data which google also has plenty of. Major advancements in AI will come from less glamorous things like labeled data available for models to learn from.

The code without the data is useless. The algorithms are only part of the equation here...and giving them away without the requisite data is useful, but not the critical path from profiting from AI. It's also not going to advance AI much. These AI labs publish subsets of the research they actually do (even if it is still a generous amount which is great).

We can even see this from OpenAI's gym efforts. These environments are creating fundamental infrastructure for pushing the boundaries on reinforcement learning.

That being said, research is a component of the problem, but even most "AI" startups just git clone some open source code and run something pretrained (say: opencv, kaldi for audio,..) and then wrap it in a nice gui.

The main things these startups focus on is delivery of the product just like the rest of these startups, very few are actually building novel algorithms. A lot of it is just them collecting data.

That being said, this is also why a lot of the novel research happens in the major for profit ad tech companies. They have the data to do research on and they can choose what to publish and what to profit from in products.

Disclosure: I work at an AI startup.

Re: Marc Andreessen on the atomization of AI

#24
post #19

I was thinking about this notion of using simulated worlds for doing training. An open source simulated world would be fantastic - to the extent that it is actually realistic. For example, maybe collect archived gameplay from some MMORPG and use it to build the simulated world. As new players play in the game, the gameplay data is collected and available for the community to use. Obviously this is not well thought ou…

https://gym.openai.com Just needs integration with popular game engines like UE and unity3d, at the moment it mostly works with 8bit game emulators

Fantastic. We should find a way to place these learning environments at random locations on a world map and interact with them via the popular VR devices :-)

Re: Marc Andreessen on the atomization of AI

#25

The common saying is "more data usually beats a better algorithm". However, this is only for short term progress. Long term progress in the field of AI clearly requires better algorithms, and doing more with less data is exactly the kind of problem that a startup in the field could solve with a clever idea. That said, these startups will have to be extremely research heavy — how does one draw the most talented resear…

>> doing more with less data is exactly the kind of problem that a startup in the field could solve with a clever idea. It will take much more than a "clever idea" to overcome the reliance on data of the current AI state of the art- that is to say, machine learning. If you think about it, machine learning algorithms are essentially clever search procedures for some optimum in a heap of data (that's optimisation algor…

There's a lot of ongoing work in machine learning focusing on learning from few examples, online learning, more biologically plausible learning and so on. Most of it is done in the academia, and some of it is done at large corporations.

Startups generally have a short runway before they run out of cash, which isn't long enough to sustain this kind of exploratory work.

Re: Marc Andreessen on the atomization of AI

#26
post #5

The common saying is "more data usually beats a better algorithm". However, this is only for short term progress. Long term progress in the field of AI clearly requires better algorithms, and doing more with less data is exactly the kind of problem that a startup in the field could solve with a clever idea. That said, these startups will have to be extremely research heavy — how does one draw the most talented resear…

Vanilla deep learning is so far ahead of what most companies use for analytics, that very little research is needed to offer a product with significant advantages over what's in practice now. In fact, I'd say most ML/DL/AI startups don't need to do research at all. They just need to spread more widely what exists now. I disagree with Andreessen that Google's Tensorflow propaganda on Udacity is "opening the kimono" or…

> Vanilla deep learning is so far ahead of what most companies use for analytics, that very little research is needed to offer a product with significant advantages over what's in practice now

But by "vanilla deep learning", I think you are referring to this vs research into new models and algorithms. I think "research" in the context of "research-heavy start-ups" is not about new algorithms, but about finding ways to use even "vanilla deep learning" on their available datasets. Quite a lot of papers at conferences, for example, are not about new algorithms, but about new ways to use old algorithms.

Re: Marc Andreessen on the atomization of AI

#27

The common saying is "more data usually beats a better algorithm". However, this is only for short term progress. Long term progress in the field of AI clearly requires better algorithms, and doing more with less data is exactly the kind of problem that a startup in the field could solve with a clever idea. That said, these startups will have to be extremely research heavy — how does one draw the most talented resear…

Truly doing much with little data may simply be impossible; its not a magical black box after all, but just as ideal a learner as possible - but if small data cannot give it confidence to confirm the existence of some subtile pattern, on what grounds is even an ideal learner supposed to believe in it?

For example, no smarts would ever confirm the Higgs in the tevatron in its 10 year run even though the machine produced thousands of Higges; its a subtle signal and you just don't have the sample size to be more than 3 sigma sure there's something there (nor where exactly).

What is key, however, is being able to leverage more than just explicit training sets, like with transfer learning and unsupervised learning, as well as making algorithms scale well with more data. The latter seems promising with deep learning, the former still a research thing.

"more data beats better algorithm" is obviously false when the algorithms don't scale with data as well, as deep learning has convincingly demonstrated, being just as mediocre as anything else on small datasets, but being leaps above everything else on sufficiently large ones.

Re: Marc Andreessen on the atomization of AI

#28
post #9

Earlier quoted context omitted.

Have you heard of the OpenAI initiative?

Yes, do you know of the number of groups who are actively developing AGI? who have made steady progress over the years? I see no big names invested in them. Why would someone try to create your own group having no background in development of AGI? Seems a bit off the mark no? Most AGI focused ventures are not even in the U.S. A prominent one is in China. Another is in Europe. The ones in the U.S are privately funded…

>do you know of the number of groups who are actively developing AGI? who have made steady progress over the years?

I do not. It seems like researchers who call themselves AGI researchers have made no more significant progress than researchers developing specific analysis techniques. Do you have a list of these groups? I would be very interested to read about their approach and progress.

On the human brain initiative it seems like their is such a huge gap between AGI and current techniques that these HBI projects (and the US response[0]) are a bit misguided into giant neural simulations.

> ... companies specifically pursuing AGI and I see a whole range of individuals ...

I completely agree this is the necessary approach to pursuing AGI. Who are these companies? They sound awesome and I would like to learn more. Are you talking OpenAI? Some other companies?

Geohot. Will check him out.

EDIT: oh, George Hotz. He is using deep learning[1]

Agreed deep learning will be a tool, but not the end of AGI. Companies padding with deep learning PhDs are looking more at specific tasks that can be tackled with deep learning. Also, I would argue that deep learning has become an umbrella term for all neural network based approaches (great marketing) and there are still great advances to be made with DL building blocks.

[0] https://en.m.wikipedia.org/wiki/BRAIN_Initiative

[1] http://www.bloomberg.com/features/2015-george-hotz-self-driv...

Re: Marc Andreessen on the atomization of AI

#29

The common saying is "more data usually beats a better algorithm". However, this is only for short term progress. Long term progress in the field of AI clearly requires better algorithms, and doing more with less data is exactly the kind of problem that a startup in the field could solve with a clever idea. That said, these startups will have to be extremely research heavy — how does one draw the most talented resear…

Truly doing much with little data may simply be impossible; its not a magical black box after all, but just as ideal a learner as possible - but if small data cannot give it confidence to confirm the existence of some subtile pattern, on what grounds is even an ideal learner supposed to believe in it? For example, no smarts would ever confirm the Higgs in the tevatron in its 10 year run even though the machine produc…

> Truly doing much with little data may simply be impossible; its not a magical black box after all, but just as ideal a learner as possible

This is interesting, and something that needs further explored IMO. I've been doing research on extracting mutual information from noisy, shifted copies of a ground truth signal, and it turns out there is a crossover point where recovering the ground truth essentially becomes impossible (in an information theoretic sense). What's interesting is that this crossover point is sharp and it looks a lot like a phase transition that one might see in statistical physics.

We need more research that provides limits on what we are capable of predicting from a given dataset — an upper bound, in other words. This would let us know if it's worth it to spend time trying to get more predictivity out of a dataset, or if the data just simply doesn't contain sufficient enough information.

Re: Marc Andreessen on the atomization of AI

#30
post #9

Earlier quoted context omitted.

Have you heard of the OpenAI initiative?

Yes, do you know of the number of groups who are actively developing AGI? who have made steady progress over the years? I see no big names invested in them. Why would someone try to create your own group having no background in development of AGI? Seems a bit off the mark no? Most AGI focused ventures are not even in the U.S. A prominent one is in China. Another is in Europe. The ones in the U.S are privately funded…

[deleted]
Post reply on HN