Live data from Hacker News

We accidentally solved robotics by watching 1M hours of YouTube

ksagar.bearblog.dev

121–130 of 183 posts

Re: We accidentally solved robotics by watching 1M hours of YouTube

#121

Earlier quoted context omitted.

> Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. I'm not sure that's necessarily true for a lot of tasks. A good way to measure this in your head is this: "If you were given remote control of two robot arms, and just one camera to look through, how many different tasks do you think you could comp…

Humans have innate knowledge that help them interact with the world and can learn from physical interaction for the rest. RGB images aren't enough.

Video games have shown that we can control pretty darn well characters in virtual worlds where we have not experienced their physics. We just look at a 2D monitor and using a joystick/keyboard we manage to figure it out.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#122
post #52

I just wrote a reply to a comment talking about the AI tells this writing has, but it got flagged so my comment disappeared when I hit post. I'll rephrase out of spite: My first thought upon reading this was that an LLM had been instructed to add a pithy meme joke to each paragraph. They don't make sense in context, and while some terminally online people do speak in memes, those people aren't quoting doge in 2025. T…

I don’t know, 400k people are listening to the White House streaming lo-fi hip hop on X right now with cutesy videos of Trump on one side and his executive orders streaming on the other at 4am. I think there’s plenty of people quoting doge in 2025.

If you’re in the US, you likely work with them and they have learned to studiously avoid talking about politics except in vagaries to avoid conflict.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#123
post #52

I just wrote a reply to a comment talking about the AI tells this writing has, but it got flagged so my comment disappeared when I hit post. I'll rephrase out of spite: My first thought upon reading this was that an LLM had been instructed to add a pithy meme joke to each paragraph. They don't make sense in context, and while some terminally online people do speak in memes, those people aren't quoting doge in 2025. T…

>those people aren't quoting doge in 2025 Could you explain what this means? Is this article quoting doge?

There was a clear attempt at the doge meme format, yes:

> very scientific. much engineering.

Emphasis on attempt because you're supposed to use words with grammatically incorrect modifiers, and the first one doesn't. (Even the second one doesn't seem entirely incorrect to me? I'm not a native speaker though.) "many scientific, so engineering" for example would have worked.

I assume they, or most likely their LLM, tried too hard to follow the most popular sequence (very, much, wow) and failed at it.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#125
post #97
post #93

Earlier quoted context omitted.

Not conservative but I used to love the meme before it was co-opted by musk, so I will occasionally use it as a "haha now you feel OLD" without thinking of its modern connotations.

Also I think it’s somehow important to not let fascism steal our cultural heritage, even if it’s just a meme. In my country, far righters are displaying the country’s flag everywhere. Now you can’t display a French flag without being thought as a far right person. That’s honestly insufferable. I know it’s less important with doge but still : before being a crypto it was just a picture of an overly innocent and enthus…

Lets see you try to recover the swastika from fascism ;)

Re: We accidentally solved robotics by watching 1M hours of YouTube

#126
post #78

Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. For example, so that you don't crush a human when doing massage (but still need to press hard), or apply the right amount of force (and finesse?) to skin a fish fillet without cutting the skin itself. Practically in the near term, it's hard to sample…

> Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. I'm not sure that's necessarily true for a lot of tasks. A good way to measure this in your head is this: "If you were given remote control of two robot arms, and just one camera to look through, how many different tasks do you think you could comp…

Humans did not accumulate that intuition just using images. In the example you gave, you subconsciously augment the image information with a lifetime of interacting with the world using all the other senses.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#127
post #109

Earlier quoted context omitted.

> Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. I'm not sure that's necessarily true for a lot of tasks. A good way to measure this in your head is this: "If you were given remote control of two robot arms, and just one camera to look through, how many different tasks do you think you could comp…

I think you vastly underestimate how difficult the task you are proposing would be without depth or pressure indication, even for a super intelligence like humans. Simple concept, pick up a glass and pour its content into a vertical hole the approximate size of your mouth. Think of all the failure modes that can be triggered in the trivial example you do multiple times a day, to do the same from a single camera feed…

The point is how much non-vision sensors vs pure vision, helps humans to be humans. Don't you think this point was proven by LLMs already that generalizability doesn't come from multi-modality but by scaling a single modality itself? And jepa is for sure designed to do a better job at that than an LLM. So no doubt about raw scaling + RL boost will kick-in highly predictable & specific robotic movements.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#128

Earlier quoted context omitted.

>those people aren't quoting doge in 2025 Could you explain what this means? Is this article quoting doge?

There was a clear attempt at the doge meme format, yes: > very scientific. much engineering. Emphasis on attempt because you're supposed to use words with grammatically incorrect modifiers, and the first one doesn't. (Even the second one doesn't seem entirely incorrect to me? I'm not a native speaker though.) "many scientific, so engineering" for example would have worked. I assume they, or most likely their LLM, tri…

"Much engineering was required" Archaic but still used a bit in articles or to give a certain vibe.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#129
>camera pose sensitivity

>the model is basically a diva about camera positioning. move the camera 10 degrees and suddenly it thinks left is right and up is down.

Reminds that several years ago Tesla had to finally start to explicitly extract 3D model from the net. Similarly i expect here that it would get pipelined - one model extracts/builds 3D, and the other is actually the "robot" working in that 3D. Each one can be alone trained much better and efficiently, with much better transfer and generalization, than the large monolithic model working from the 2D video. In pipeline approach, it is very easy to generate synthetic input 3D data better covering interesting scenario space for the "robot" model.

And, for example, you can't just, without significant training, feed the large monolithic model a lidar point space instead of videos. Whereis in a pipelined approach, you just switch the 3D generating pipeline input model.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#130

Earlier quoted context omitted.

Humans have innate knowledge that help them interact with the world and can learn from physical interaction for the rest. RGB images aren't enough.

Video games have shown that we can control pretty darn well characters in virtual worlds where we have not experienced their physics. We just look at a 2D monitor and using a joystick/keyboard we manage to figure it out.

Yeah but we already have a conception of what physics should be prior to that that helps us enormously. It's not like game designers are coming up with stuff that intentionally breaks our naïve physics.
Post reply on HN