Live data from Hacker News

Gemini Robotics 2 brings whole body intelligence to robots

deepmind.google

561–570 of 583 posts

Re: Gemini Robotics 2 brings whole body intelligence to robots

#561

Earlier quoted context omitted.

No, the above is absolutely exactly true. Robots folding clothes and tying shoelaces etc is nothing but a tech demo at this point. It's difficult to grok this because if you watch a human folding a t-shirt, you can reliably predict that the same human will fold a different t-shirt just as well, and in fact be perfectly capable of folding a wide variety of other clothes items as well. Not so for robots. With robots, w…

You make some very good points about the ACT-2 if it’s nine garment types that’s still very acceptable but the point you bought up about them training on exactly those pieces of garments is a possibility.

Cheers. You can see that's what they did if you look at the 4th video on the page, under the figure titled "Quality Remains High Across Garment Types" right below the paragraph that starts with "To put these scores in context". Sorry, I have no idea how to link to that video specifically.

In the left half of the video you can see that the robot is (trying to) exactly match the folds of the human in the right half and it's doing so while folding the exact same garments on the exact same surface.

The right half of the video is not a training demonstration, I don't think, since the robot is trained by teleoperation AFAICT (it needs to because it must use its head-mounted camera to control its movements) but that just underlines the degree to which their training regime is exactly copying the movements of a trainer, on the same garment, in the same environment.

This is a limitation of the training approach, by RL. With RL you learn a mapping between sets of pixels (as in the video that comes in through the robot's camera) and robot actions (as in actuator commands). What that means is that once a policy is trained and the robot is deployed, if the input pixels are significantly different than the input pixels at training the robot doesn't have a policy that matches the input pixels and so it can't find the right actions to take. So they have to keep the training and deployment garments and even the environments the same, or as similar as possible.

You can see some more evidence of this in the video right under the paragraph with the title "Hill-Climbing Reliability Through Post-Training". The robot at the front of the video, with the bright red trim, is shown trying to fold a grey t-shirt with white flower decorations and a frilly hem (how adorable before the robot has completed the fold. If it could complete it, you can rest assured that the video would be showing off the entire folding sequence as it does for the robot with the green cap on the other side of the bed.

That paragraph is making a claim about a "post-training" regime that's supposed to improve generalisation but it leaves more details to a "separate technical post". So I can't tell what it's supposed to be doing, but I don't think it's working.

When I watch videos like that I always remind myself that a) I'm watching a tech demo created to attract investment and b) I've watched way too many of those, going all the way back to the Boston Dynamic videos of Robot Dog or of Atlas doing backflips and yet the state of the art hasn't really budged since. Such videos make it easy to overestimate the state of the art in autonomous robotics and in fact are meant do precisely that: play up robots' true capabilities. It's just impossible to say anything about a robot's general capabilities by watching a few minutes or even a few hours of video. OtoH if you know what to look for you can tell everyone is basically stuck at the same level and trying the same things to escape it. The truth is robotic autonomy is several major breakthroughs away and nobody has any idea how to get there. So we'll be seeing many more of those tech demo videos in the years to come.

Re: Gemini Robotics 2 brings whole body intelligence to robots

#562
post #560

Earlier quoted context omitted.

Depends entirely on VLA arch. Some have dedicated action diffusion heads that work in a standalone non-text action output space. Much like an LLM can either use an external TTS or have audio output heads attached to it directly for native S2S. But your entire premise is wrong regardless of that. Even if VLAs were forever bound to outputting text, you'd have to prove that they're fundamentally incapable of emitting te…

>But your entire premise is wrong regardless of that. You don't understand what I am saying. The crux of your misunderstanding is here >emitting text that maps to useful action sequences If you have a static mapping from text to action, then you are throwing away all the advantage of using an AI. The whole point of AI is that you can get an output from an input without explicit mapping. So If you use explicit mapping…

You can literally have an LLM output target joint angles. As text. To be decoded by an explicit decoder, and executed by the robot. Some early VLAs did exactly that.

Your entire premise is wrong.

Modern action decoders are different, and usually take the form of neural networks trained end to end jointly with the rest of the model. Not fundamentally more expressive, just more in line with what we want.

Re: Gemini Robotics 2 brings whole body intelligence to robots

#563
post #560

Earlier quoted context omitted.

>But your entire premise is wrong regardless of that. You don't understand what I am saying. The crux of your misunderstanding is here >emitting text that maps to useful action sequences If you have a static mapping from text to action, then you are throwing away all the advantage of using an AI. The whole point of AI is that you can get an output from an input without explicit mapping. So If you use explicit mapping…

You can literally have an LLM output target joint angles. As text. To be decoded by an explicit decoder, and executed by the robot. Some early VLAs did exactly that. Your entire premise is wrong. Modern action decoders are different, and usually take the form of neural networks trained end to end jointly with the rest of the model. Not fundamentally more expressive, just more in line with what we want.

You are repeating "It can be done, trust me!". But that is not very convincing.

Re: Gemini Robotics 2 brings whole body intelligence to robots

#564
post #563

Earlier quoted context omitted.

You can literally have an LLM output target joint angles. As text. To be decoded by an explicit decoder, and executed by the robot. Some early VLAs did exactly that. Your entire premise is wrong. Modern action decoders are different, and usually take the form of neural networks trained end to end jointly with the rest of the model. Not fundamentally more expressive, just more in line with what we want.

You are repeating "It can be done, trust me!". But that is not very convincing.

I'm repeating "what you claim to be impossible was done 3 years ago and was already replaced with better versions of the same idea and you are hilariously out of touch".

Re: Gemini Robotics 2 brings whole body intelligence to robots

#565
post #563

Earlier quoted context omitted.

You are repeating "It can be done, trust me!". But that is not very convincing.

I'm repeating "what you claim to be impossible was done 3 years ago and was already replaced with better versions of the same idea and you are hilariously out of touch".

Show me one video where a robot follows textual prompts and come up with its own movements to solve the prompt.

Re: Gemini Robotics 2 brings whole body intelligence to robots

#566
post #268

Earlier quoted context omitted.

If that ends up being true, they should be very surprised. We're looking at ~$725B combined hyperscaler capex in 2026 (on a path to $1.08T by 2028) against roughly $25B of AI service revenue in 2025 on $250B+ of infrastructure spend. By 2030, the global data center build-out will require $6.7 trillion in capital expenditure. If hyperscalers require a 25% return on AI-specific capex, the industry needs to generate ~$1…

You know SpaceX is still publicly valued at $1.5T marketcap, right? (offering price has nothing to do with anything other than ego of founders/bankers). Yes, I will wait for the correct valuation when all stocks enter the liquidity pool. OpenAI/Anthropic obscenely profitable is an easy bet (for me)

Better you holding the bag than me. Unfortunately all of us will be holding parts of it.

Re: Gemini Robotics 2 brings whole body intelligence to robots

#567
post #565

Earlier quoted context omitted.

I'm repeating "what you claim to be impossible was done 3 years ago and was already replaced with better versions of the same idea and you are hilariously out of touch".

Show me one video where a robot follows textual prompts and come up with its own movements to solve the prompt.

That's just about any video of any VLA ever. Including Gemini Robotics 2.

Re: Gemini Robotics 2 brings whole body intelligence to robots

#568
post #533

Earlier quoted context omitted.

Exactly, the last year or two has been insanely innovative for actuators, with some new material combinations, approaches and multiple startups that are starting to scale up their innovations in this space. Give them a few years to sort out the vaporwares and the robotics revolution will begin.

What innovations specifically?

Can't really name them off the top of my head, sorry!

But if you look up on Google or X, there are multiple companies doing things like pressure based actuators (Clone robotics) or mesh-based knit ones, new forms of EFAM, some new soft actuators and more.

Re: Gemini Robotics 2 brings whole body intelligence to robots

#569
post #565

Earlier quoted context omitted.

Show me one video where a robot follows textual prompts and come up with its own movements to solve the prompt.

That's just about any video of any VLA ever. Including Gemini Robotics 2.

Should be trivial to link to one then...

Re: Gemini Robotics 2 brings whole body intelligence to robots

#570

Earlier quoted context omitted.

you mean like the self driving car which elon promised ten years ago. :D Robotic vision AI is far more complex than an LLM

Yeah, kinda. Waymos have been on the road since 2023.

Oh come on, level 4 / Waymo is a joke. That’s a good weather system…everything under level 5 is not really self driving cause you can’t really rely on it. It’s just a nice hobby research project of some billionaires :D.

At self driving system I’m think about bmw which drives automatically at 230km/h on a German highway and races at aggressive as I am at the high way cause I don’t want to be more slow that the train at the distance from Berlin to Munich ;)

Post reply on HN