Live data from Hacker News

We accidentally solved robotics by watching 1M hours of YouTube

ksagar.bearblog.dev

151–160 of 183 posts

Re: We accidentally solved robotics by watching 1M hours of YouTube

#151

Earlier quoted context omitted.

> Who cares at this point Anyone who has a shred of integrity. I'm not a fan of overreaching copyright laws, but they've been strictly enforced for years now. Decades, even. They've ruined many lives, like how they killed Aaron Swartz. But now, suddenly, violating copyright is totally okay and carries no consequences whatsoever because the billionaires decided that's how they can get richer now? If you want to even t…

Aaron Swartz died of suicide, not copyright. His death was a tragedy but it wasn't done to him.

Crimes generally don't kill the criminal. It's the reaction by authorities that kills (perceived) criminals.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#152

I didn't understand a single word about this post and what was supposed to be solved and had to stop reading. Was this actually written by a human being? If so, the author(s) suffer from severe language communication problems. Doesn't seem to be grounded at least with reality and my personal experience with robotics. But here's my real world take: Robotics is going to be partially solved when ROS/ROS2 becomes effecti…

It's not a scholarly article but a blog post but you're still right to be frustrated at the very bad writing. I do get the jargon, despite myself, so I can translate: the authors of the blog post claim that machine learning for autonomous robotics is "solved" thanks to an instance of V-JEPA 2 trained on all videos on youtube. It isn't, of course, and the authors themselves point out the severe limitations of the otherwise promising approach (championed by Yan LeCun) when they say, in a notably more subdued manner:

>> the model is basically a diva about camera positioning. move the camera 10 degrees and suddenly it thinks left is right and up is down.

>> in practice, this means you have to manually fiddle with camera positions until you find the sweet spot. very scientific. much engineering.

>> long-horizon drift

>> try to plan more than a few steps ahead and the model starts hallucinating.

That is to say, not quite ready for the real world, V-JEPA 2 is.

But for those who don't get the jargon there's a scholarly article linked at the end of the post that is rather more sober and down-to-earth:

A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-scale video data with a small amount of interaction data (robot trajectories), to develop models capable of understanding, predicting, and planning in the physical world. We first pre-train an action-free joint-embedding-predictive architecture, V-JEPA 2, on a video and image dataset comprising over 1 million hours of internet video. V-JEPA 2 achieves strong performance on motion understanding (77.3 top-1 accuracy on Something-Something v2) and state-of-the-art performance on human action anticipation (39.7 recall-at-5 on Epic-Kitchens-100) surpassing previous task-specific models. Additionally, after aligning V-JEPA 2 with a large language model, we demonstrate state-of-the-art performance on multiple video question-answering tasks at the 8 billion parameter scale (e.g., 84.0 on PerceptionTest, 76.9 on TempCompass). Finally, we show how self-supervised learning can be applied to robotic planning tasks by post-training a latent action-conditioned world model, V-JEPA 2-AC, using less than 62 hours of unlabeled robot videos from the Droid dataset. We deploy V-JEPA 2-AC zero-shot on Franka arms in two different labs and enable picking and placing of objects using planning with image goals. Notably, this is achieved without collecting any data from the robots in these environments, and without any task-specific training or reward. This work demonstrates how self-supervised learning from web-scale data and a small amount of robot interaction data can yield a world model capable of planning in the physical world.

https://arxiv.org/abs/2506.09985

In other words, some interesting results, some new SOTA, some incremental work. But lots of work for a big team of a couple dozen researchers so there's good stuff in there almost inevitably.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#153
post #14

This was a bit hard to read. It would be good to have a narrative structure and more clear explanation of concepts.

Very intentional. Their response would be: “if you need narrative structure and clear explanation of concepts, yngmi”.

And the answer to that would be: WNGTI.

https://www.youtube.com/watch?v=4xmckWVPRaI

Capitalia tantum.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#154

Solved??? Where?

Yeah, wake me up when they have a robot that can wash, peel, cut fruit and vegetables; unwrap, cut, cook meat; measure salt and spices; whip cream; knead and shape dough; and clean up the resulting mess from all of these. Then they will have "solved" part of robotics.

>> Yeah, wake me up when they have a robot that can wash, peel, cut fruit and vegetables; unwrap, cut, cook meat; measure salt and spices; whip cream; knead and shape dough; and clean up the resulting mess from all of these.

Someone's getting peckish :P

Re: We accidentally solved robotics by watching 1M hours of YouTube

#155
post #52

I just wrote a reply to a comment talking about the AI tells this writing has, but it got flagged so my comment disappeared when I hit post. I'll rephrase out of spite: My first thought upon reading this was that an LLM had been instructed to add a pithy meme joke to each paragraph. They don't make sense in context, and while some terminally online people do speak in memes, those people aren't quoting doge in 2025. T…

>> They don't make sense in context, and while some terminally online people do speak in memes, those people aren't quoting doge in 2025.

Cringely, they are. Nobody who isn't desperate to appear cool would write in that terminally grating register, including when using an LLM to do the writing.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#156

Earlier quoted context omitted.

> some terminally online people do speak in memes, those people aren't quoting doge in 2025. You may be surprised to find out how incorrect this. I can think of two popular conservative sites likely to quote Doge people off hand that do this. I read all news in order not be an insufferable ideologue. So again, off the top of my head, NotTheBee (I think affiliated to BabylonBee (conservative The Onion)) and Twitchy. A…

They are referring to the original doge meme of the dog, not the government initiative today. I guess "quote" isn't really the right word, more like "doing"

A reoccurring mistake in this thread. I blame Elon Musk and his boomer humer.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#157
post #107

Earlier quoted context omitted.

Good catch. Approximately 9,192,631 Turkish decibels.

Fun fact: the International Bureau of Weights and Measures in Paris is the owner of a perfect 0 dB noise floor enclosed in a perfect titanium sphere (with some sheep's wool filling to avoid reflections). There is a small door on the side over which microphone capsules can be inserted for calibration. (/joke)

Too bad the joke doesn't work if you understand decibels.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#158
post #2

Does YouTube allow massive scraping like this in their ToS?

Per HiQ vs. LinkedIn, it doesn't matter what their ToS says if the scraper didn't have to agree to the ToS to scrape the data. YouTube will serve videos to someone who isn't logged in. So if you've never agreed to YouTube's ToS, you can scrape the videos. If YT forced everyone to log in before they could watch a video, then anyone who wants to scrape videos would have had to agree to the ToS at some point.

It won't serve me videos if I'm not logged in. It tells me to sign in to prove I'm not a bot. How do these people get around this?

Re: We accidentally solved robotics by watching 1M hours of YouTube

#159
post #78

Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. For example, so that you don't crush a human when doing massage (but still need to press hard), or apply the right amount of force (and finesse?) to skin a fish fillet without cutting the skin itself. Practically in the near term, it's hard to sample…

  > Pure vision will never be enough because it does not contain information
Say it louder for those in the back!

But actually there's more to this that makes the problem even harder! Lack of sensors is just the beginning. There's well known results in physics that:

  You cannot create causal models through observation alone.
This is a real pain point for these vision world models and most people I talk to (including a lot at the recent CVPR) just brush this off as "we're just care if it works." Guess what?! Everyone that is pointing this out also cares that it works! We need to stop these thought terminating cliches. We're fucking scientists.

Okay, so why isn't observation enough? It's because you can't differentiate alternative but valid hypotheses. You often have to intervene! We're all familiar with this part. You control variables and modify one or a limited set at a time. Experimental physics is no easy task, even for things that sound rather mundane. This is in fact why children and animals play (okay, I'm conjecturing here).

We need to mention chaos here, because it's the easiest way to understand this. There's many famous problems that fall into this category like the double pendulum, 3 Body Problem, or just fucking gas molecules moving around. Let's take the last one. Suppose you are observing some gas molecules moving inside a box. You measure their positions at t0 and at T. Can you predict their trajectories between those time points? Surprisingly, the answer is no. You can only do this statistically. There's probably paths but not deterministic (this same logic is what leads to multiverse theory btw). But now suppose I was watching the molecules too, but I was continuously recording between t0 and T. Can I predict the trajectories? Well, I don't need to, I just write it down.

Now I hear you, you're saying "Godelski, you observed!" But the problem with these set of problems is that if you don't observe the initial state you can't predict moving forwards and if you don't have very precise observation intervals you are hit with the same problem. I you turn around while I start a double pendulum you can have as much time as you want when you turn back around, you won't be able to model its trajectories.

But it gets worse still. There are confounding variables. There is coupling. Difficult to differentiate hypotheses via causal ordering. And so so much more. If you ever wonder why physicists do so much math it's because doing that is a fuck ton easier than doing the whole set of testing and then reverse engineering the equations from those observations. But in physics we care about counterfactual statements. In F=ma we can propose new masses and new accelerations and rederive the results. That's the what it is all about. Your brain does an amazing job at this too! You need counterfactual modeling to operate in real world environments. You have to be able to ask and answer "what happens if that kid runs into the street?"

I highly suggest people read The Relativity of Wrong [0]. Its a short essay by Isaac Asimov that can serve as a decent intro, though far from complete. I'm suggesting it because I don't want people to confuse "need counterfactual model" with "need the right answer." If you don't get into metaphysics, these results will be baffling.[1] It is also needed to answer any confusion you might have around the aforementioned distinction.

Tldr:

  if you could do it from observation alone, physics would have been solved a thousand years ago
There's a lot of complexity and depth that is easy to miss with the excitement, but it still matters.

I'm just touching the surface here too, and we're just talking about mechanics. No quantum needed, just information loss

[0] https://hermiene.net/essays-trans/relativity_of_wrong.html

[1] maybe this is why there are so few physicists working on the world modeling side of ML. At least, using that phrase...

Re: We accidentally solved robotics by watching 1M hours of YouTube

#160
post #78

Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. For example, so that you don't crush a human when doing massage (but still need to press hard), or apply the right amount of force (and finesse?) to skin a fish fillet without cutting the skin itself. Practically in the near term, it's hard to sample…

> Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. I'm not sure that's necessarily true for a lot of tasks. A good way to measure this in your head is this: "If you were given remote control of two robot arms, and just one camera to look through, how many different tasks do you think you could comp…

  > because you as a human have really good intuition about the world.
This is the line that causes your logic to fail.

You introduced knowledge not obtained through observation. In fact, the knowledge you introduced is the whole chimichanga! It is an easy mistake to make, so don't feel embarrassed.

The claim is that one can learn a world model[0] through vision. The patent countered by saying "vision is not enough." Then you countered by saying "vision is enough if you already have a world model."

[0] I'll be more precise here. You can learn *A* world model, but it isn't the one we really care about and "a world" doesn't require being a self consistent world. We could say the same thing about "a physics", but let's be real, when we say "physics" we know which one is being discussed...

Post reply on HN