Live data from Hacker News

We accidentally solved robotics by watching 1M hours of YouTube

ksagar.bearblog.dev

111–120 of 183 posts

Re: We accidentally solved robotics by watching 1M hours of YouTube

#111
I do not know and do not care much about robotics per se, but I wish LLM's were better with spatial reasoning. If the new insight helps with that - great!

I dabbled a bit in geolocation with LLM's recently. It is still surprising to me how good they are with finding the general area a picture was taken. Give it a photo of a random street corner on this earth and it is likely will not only tell you the correct city or town but most often even the correct quarter.

On the other hand, if you ask it for a birds eye view of a green, a brown and a white house on the north side of a one-way street (running west to east) east of an intersection running north to south, it may or may not get it right. If you want it to add an arrow going in the direction of the one-way street, it certainly has no clue at all and the result is 50/50.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#112
post #14

This was a bit hard to read. It would be good to have a narrative structure and more clear explanation of concepts.

It would also be good if the perspective of the article would stay put. This "we" and "they" thing was at best confusing and at worst possibly a way to get more clicks or pretend the author had something to do with the work.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#113
post #68

Earlier quoted context omitted.

What's the point in writing something while "not caring" if the reader understands or not? Seems like a false confidence or false bravado to me; it reads like an attempt to project an impression, and not really an attempt to communicate.

Basically: If you understand the topic well, you’re not the target audience. This is a type of information arbitrage where someone samples something intellectual without fully understanding it, then writes about it for a less technical audience. Their goal is to appear to be the expert on the topic, which translates into clout, social media follows, and eventually they hope job opportunities. The primary goal of the…

I guess "bullshitting as a career" isn't going away any time soon.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#114
post #109

Earlier quoted context omitted.

> Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. I'm not sure that's necessarily true for a lot of tasks. A good way to measure this in your head is this: "If you were given remote control of two robot arms, and just one camera to look through, how many different tasks do you think you could comp…

I think you vastly underestimate how difficult the task you are proposing would be without depth or pressure indication, even for a super intelligence like humans. Simple concept, pick up a glass and pour its content into a vertical hole the approximate size of your mouth. Think of all the failure modes that can be triggered in the trivial example you do multiple times a day, to do the same from a single camera feed…

If I have to pour water into my mouth, you can bet it's going all over my shirt. That's not how we drink.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#115
post #109

Earlier quoted context omitted.

> Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. I'm not sure that's necessarily true for a lot of tasks. A good way to measure this in your head is this: "If you were given remote control of two robot arms, and just one camera to look through, how many different tasks do you think you could comp…

I think you vastly underestimate how difficult the task you are proposing would be without depth or pressure indication, even for a super intelligence like humans. Simple concept, pick up a glass and pour its content into a vertical hole the approximate size of your mouth. Think of all the failure modes that can be triggered in the trivial example you do multiple times a day, to do the same from a single camera feed…

A routine gesture I've done everyday for almost all my life: getting a glass out of the shelves and into my left hand. It seems like a no brainer, I open the cabinet with my left hand, take the glass with my right hand, throw the glass from my right hand to the left hand while closing the cabinet with my shoulder. Put the glass under the faucet with left hand, open the faucet with the right hand.

I have done this 3 seconds gesture, and variations of it, my whole life basically, and never noticed I was throwing the glass from one hand to the other without any visual feedback.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#116
post #78

Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. For example, so that you don't crush a human when doing massage (but still need to press hard), or apply the right amount of force (and finesse?) to skin a fish fillet without cutting the skin itself. Practically in the near term, it's hard to sample…

> Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. I'm not sure that's necessarily true for a lot of tasks. A good way to measure this in your head is this: "If you were given remote control of two robot arms, and just one camera to look through, how many different tasks do you think you could comp…

> When you start thinking about it, you realize there are a lot of things you could do with just the arms and one camera, because you as a human have really good intuition about the world.

And where does this intuition come from? It was buily by also feeling other sensations in addition to vision. You learned how gravity pulls things down when you were a kid. How hot/cold feels, how hard/soft feels, how thing smell. Your mental model of the world is substantially informed by non-visual clues.

> It therefore follows that robots should be able to learn with just RGB images too!

That does not follow at all! It's not how you learned either.

Neither have you learned to think by consuming the entirety of all text produced on the internet. LLMs therefore don't think, they are just pretty good at faking the appearance of thinking.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#118
post #52

I just wrote a reply to a comment talking about the AI tells this writing has, but it got flagged so my comment disappeared when I hit post. I'll rephrase out of spite: My first thought upon reading this was that an LLM had been instructed to add a pithy meme joke to each paragraph. They don't make sense in context, and while some terminally online people do speak in memes, those people aren't quoting doge in 2025. T…

>those people aren't quoting doge in 2025

Could you explain what this means? Is this article quoting doge?

Re: We accidentally solved robotics by watching 1M hours of YouTube

#119
post #78

Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. For example, so that you don't crush a human when doing massage (but still need to press hard), or apply the right amount of force (and finesse?) to skin a fish fillet without cutting the skin itself. Practically in the near term, it's hard to sample…

> Pure vision will never be enough because it does not contain information about the physical feedback like pressure and touch, or the strength required to perform a task. I'm not sure that's necessarily true for a lot of tasks. A good way to measure this in your head is this: "If you were given remote control of two robot arms, and just one camera to look through, how many different tasks do you think you could comp…

Humans have innate knowledge that help them interact with the world and can learn from physical interaction for the rest. RGB images aren't enough.

Re: We accidentally solved robotics by watching 1M hours of YouTube

#120
post #109

Earlier quoted context omitted.

I think you vastly underestimate how difficult the task you are proposing would be without depth or pressure indication, even for a super intelligence like humans. Simple concept, pick up a glass and pour its content into a vertical hole the approximate size of your mouth. Think of all the failure modes that can be triggered in the trivial example you do multiple times a day, to do the same from a single camera feed…

If I have to pour water into my mouth, you can bet it's going all over my shirt. That's not how we drink.

Except this is the absolutely most common thing humans do, and my argument is that that it will spill water all over but rather that it will shatter numerous glasses, knock them over etc all before it has picked up the glass.

The same process will be repeated many times trying to move the glass to its “face” and then when either variable changes, plastic vs glass, size, shape, location and all bets are off purely because there just plainly is the enough information

Post reply on HN