Live data from Hacker News

We accidentally solved robotics by watching 1M hours of YouTube

ksagar.bearblog.dev

41–50 of 183 posts

Re: We accidentally solved robotics by watching 1M hours of YouTube

#41
Extremely oversold article.

> the core insight: predict in representation space, not pixels

We've been doing this since 2014? Not only that, others have been doing it at a similar scale. e.g. Nvidia's world foundation models (although those are generative).

> zero-shot generalization (aka the money shot)

This is easily beaten by flow-matching imitation learning models like what Pi has.

> accidentally solved robotics

They're doing 65% success on very simple tasks.

The research is good. This article however misses a lot of other work in the literature. I would recommend you don't read it as an authoritative source.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#45

This is interesting for generalized problems ( "make me a sandwich" ) but not useful for most real world functions ( "perform x within y space at z cost/speed" ). I think the number of people on the humanoid bandwagon trying to implement generalized applications is staggering right now. The physics tells you they will never be as fast as purpose-built devices, nor as small, nor as cheap. That's not to say there's zer…

Well, there’s a middle ground, kinda. Using more specialized hardware (ex: cobots) but deploy state-of-art Physical AI (ML/Computer Vision) on them. We’re building one such startup at ko-br ( https://ko-br.com/ ) :))

Quite a few startups in your space. Many deployed with customers. Good luck finding a USP!

Re: We accidentally solved robotics by watching 1M hours of YouTube

#47

This is interesting for generalized problems ( "make me a sandwich" ) but not useful for most real world functions ( "perform x within y space at z cost/speed" ). I think the number of people on the humanoid bandwagon trying to implement generalized applications is staggering right now. The physics tells you they will never be as fast as purpose-built devices, nor as small, nor as cheap. That's not to say there's zer…

Very good point! This area faces a similar misalignment of goals in that it tries to be a generic fit-all solution that is rampant with today's LLMs. We made a sandwich but it cost you 10x more than it would a human and slower might slowly become faster and more efficient but by the time you get really good at it, its simply not transferable unless the model is genuinely able to make the leap across into other domain…

Note that the username here is a Korean derogatory term for Chinese people.
Post reply on HN