Live data from Hacker News

Helix: A vision-language-action model for generalist humanoid control

figure.ai

31–40 of 178 posts

Re: Helix: A vision-language-action model for generalist humanoid control

#31
post #21
post #15

Until we get robots with really good hands, something I'd love in the interim is a system that uses _me_ as the hands. When it's time to put groceries away, I don't want to have to think about how to organize everything. Just figure out which grocery items I have, what storage I have available, come up with an optimized organization solution, then tell me where to put things, one at a time. I'm cautiously optimistic…

Maybe I don't understand exactly what you're describing but why would anyone pay for this? When I bring home the shopping I just... chuck stuff in the cupboards. I already know where it all goes. Maybe you can explain more?

Maybe they're imagining more complex tasks like working on an engine.

Re: Helix: A vision-language-action model for generalist humanoid control

#32
post #27

Does anyone know how long they have been at this? Is this mainly a reimplementation of the physical intelligence paper + the dual size/freq + the cooperative part?

"Over a year" according to the founder: https://x.com/adcock_brett/status/1892578309344502191

Re: Helix: A vision-language-action model for generalist humanoid control

#33
This is amazing but it also made me realize I just don’t trust these videos. Is it sped up? How much is preprogrammed?

I now they claim there’s no special coding but did they practice this task? Special training?

Even if this video is totally legit I’m but burned out by all the hype videos in general.

Re: Helix: A vision-language-action model for generalist humanoid control

#35
I get the impression there’s a language model sending high level commands to a control model? I wonder when we can have one multimodal model that controls everything.

The latest models seemed to be fluidly tied in with generating voice; even singing and laughing.

It seems like it would be possible to train a multimodal that can do that with low level actuator commands.

Re: Helix: A vision-language-action model for generalist humanoid control

#36
post #33

This is amazing but it also made me realize I just don’t trust these videos. Is it sped up? How much is preprogrammed? I now they claim there’s no special coding but did they practice this task? Special training? Even if this video is totally legit I’m but burned out by all the hype videos in general.

they seem slow to me, I was thinking they're slow for safety

Re: Helix: A vision-language-action model for generalist humanoid control

#38
post #33

This is amazing but it also made me realize I just don’t trust these videos. Is it sped up? How much is preprogrammed? I now they claim there’s no special coding but did they practice this task? Special training? Even if this video is totally legit I’m but burned out by all the hype videos in general.

They appear to be realtime, based on the robot's movements with the human in the scene. If you believe the article, it's zero shot (no preprogramming, practice or special training).

Re: Helix: A vision-language-action model for generalist humanoid control

#39
post #21
post #15

Until we get robots with really good hands, something I'd love in the interim is a system that uses _me_ as the hands. When it's time to put groceries away, I don't want to have to think about how to organize everything. Just figure out which grocery items I have, what storage I have available, come up with an optimized organization solution, then tell me where to put things, one at a time. I'm cautiously optimistic…

Maybe I don't understand exactly what you're describing but why would anyone pay for this? When I bring home the shopping I just... chuck stuff in the cupboards. I already know where it all goes. Maybe you can explain more?

One use case I imagine is skilled workmanship. For example, putting on a pair of AR glasses and having the equivalent of an experienced plumber telling me exactly where to look for that leak and how to fix it. Or how to replace my brake pads or install a new kitchen sink.

When I hire a plumber or a mechanic or an electrician, I'm not just paying for muscle. Most of the value these professionals bring is experience and understanding. If a video-capable AI model is able to assume that experience, then either I can do the job myself or hire some 20 year old kid at roughly minimum wage. If capabilities like this come about, it will be very disruptive, for better and for worse.

Re: Helix: A vision-language-action model for generalist humanoid control

#40
post #35

I get the impression there’s a language model sending high level commands to a control model? I wonder when we can have one multimodal model that controls everything. The latest models seemed to be fluidly tied in with generating voice; even singing and laughing. It seems like it would be possible to train a multimodal that can do that with low level actuator commands.

If you read the article, they describe a two-system approach; one "think fast" 80M parameter model running at 200hz to control motion, and one "think slow" 7B parameter model running at ~7-9hz for everything else (scene understanding, language processing, etc).

If that sounds like a cheat, neuroscientists tell us this is how the human brain works.

Post reply on HN