A fast reactive visuomotor policy that translates the latent semantic representations produced by S2 into precise continuous robot actions at 200 Hz
Why 200Hz...? Any experts in here on robotics? Because to this layman that seems really often to update motor controls.Helix: A vision-language-action model for generalist humanoid control
121–130 of 178 posts
Re: Helix: A vision-language-action model for generalist humanoid control
#122"The first time you've seen these objects" is a weird thing to say. One presumes that this is already in their training set, and that these models aren't storing a huge amount of data in their context, so what does that even mean?
It probably gives them confidence that they can accurately see a thing even though they don't know what that thing is. I could also imagine a lot of safety around leaving things outside of the current task alone so you might have to bend over backwards to get new objects worked on.
These models are trained such that the given conditions (the visual input and the text prompt) will be continued with a desirable continuation (motor function over time).
The only dimension accuracy can apply to is desirability.
Re: Helix: A vision-language-action model for generalist humanoid control
#123The demo is quite interesting but I am mostly intrigued by the claim that it is running totally local to each robot. It seems to use some agentic decision making but the article doesn't touch on that. What possible combo of model types are they stringing together? Or is this something novel? The article mentions that the system in each robot uses two ai models. S2 is built on a 7B-parameter open-source, open-weight V…
What part of this system understands 3 dimensional space of that kitchen?
The visual model "understands" it most readily, I'd say -- like a traditional Waymo CNN "understands" the 3D space of the road. I don't think they've explicitly given the models a pre-generated pointcloud of the space, if that's what you're asking. But maybe I'm misunderstanding? How does the robot closest to the refrigerator know to pass the cookies to the robot on the left?
It appears that the robot is being fed plain english instructions, just like any VLM would -- instead of the very common `text+av => text` paradigm (classifiers, perception models, etc), or the less common `text+av => av` paradigm (segmenters, art generators, etc.), this is `text+av => movements`.Feeding the robots the appropriate instructions at the appropriate time is a higher-level task than is covered by this demo, but I think is pretty clearly doable with existing AI techniques (/a loop).
How is this kind of speech to text, visual identification, decision making, motor control, multi-robot coordination and navigation of 3d space possible locally?
If your question is "where's the GPUs", their "AI" marketing page[1] pretty clearly implies that compute is offloaded, and that only images and instructions are meaningfully "on board" each robot. I could see this violating the understanding of "totally local" that you mentioned up top, but IMHO those claims are just clarifying that the individual figures aren't controlled as one robot -- even if they ultimately employ the same hardware. Each period (7Hz?) two sets of instructions are generated. What possible combo of model types are they stringing together? Or is this something novel?
Again, I don't work in robotics at all, but have spent quite a while cataloguing all the available foundational models, and I wouldn't describe anything here as "totally novel" on the model level. Certainly impressive, but not, like, a theoretical breakthrough. Would love for an expert to correct me if I'm wrong, tho!EDIT: Oh and finally:
Is anyone skeptical? How much of this is possible vs a staged tech demo to raise funding?
Surely they are downplaying the difficulties of getting this setup perfectly, and don't show us how many bad runs it took to get these flawless clips.They are seeking to raise their valuation from ~$3B to ~$40B this month, sooooooo take that as you will ;)
https://www.reuters.com/technology/artificial-intelligence/r...
Re: Helix: A vision-language-action model for generalist humanoid control
#124Earlier quoted context omitted.
do you need a robot to work in your house 24/7?
Why not? If it's done with all the chores, I can have it make some silly woodworking / art project for Etsy to earn its keep, or just loan it out to neighbors.
Re: Helix: A vision-language-action model for generalist humanoid control
#125The demo is quite interesting but I am mostly intrigued by the claim that it is running totally local to each robot. It seems to use some agentic decision making but the article doesn't touch on that. What possible combo of model types are they stringing together? Or is this something novel? The article mentions that the system in each robot uses two ai models. S2 is built on a 7B-parameter open-source, open-weight V…
I'm very far from an expert, but: What part of this system understands 3 dimensional space of that kitchen? The visual model "understands" it most readily, I'd say -- like a traditional Waymo CNN "understands" the 3D space of the road. I don't think they've explicitly given the models a pre-generated pointcloud of the space, if that's what you're asking. But maybe I'm misunderstanding? How does the robot closest to t…
their "AI" marketing page[1] pretty clearly implies that compute is offloaded
I think that answers most of my questions.I am also not in robotics, so this demo does seem quite impressive to me but I think they could have been more clear on exactly what technologies they are demonstrating. Overall still very cool.
Thanks for your reply
Re: Helix: A vision-language-action model for generalist humanoid control
#126Earlier quoted context omitted.
Imagine they bring one out to a construction site and they treat the robot as a new rookie guy, go pick up those pipes. That would be an ultimate on the fly test to me.
Picking up a bundle of loose pipes actually seems like a great benchmark for humanoid robots. Especially if they're not in a perfect pile. A full test could be something like grabbing all the pipes, from the floor, and putting them into a truck bed, in some (hopefully) sane fashion
You put a keyring with bunch of different keys in front of a robot and then instruct it pick it up and open a lock while you are describing which key is the correct one. Something like "Use the key with black plastic head and you need to put it in teeths facing down"
I have low hopes of this being possibe in the next 20 years. I hope I am still alive to witness if it ever happens.
Re: Helix: A vision-language-action model for generalist humanoid control
#127There’s nothing I want more than a robot that does house chores. That’s the real 10x multiplier for humans to do what they do best.
Yeah, except that future doesn't need us. By us I mean those of us who don't have $1B to their name. Do you really expect the oligarchs to put up with the environmental degradation of 8 billion humans when they can have a pristine planet to themselves with their whims served by the AI and these robots? I fully anticipate that when these things mature enough we'll see an "accidental" pandemic sweep and kill off 90% of…
Re: Helix: A vision-language-action model for generalist humanoid control
#128Earlier quoted context omitted.
Then you are equally fucked as the AI will be, so no difference. Case in point, I remember about ten years ago our washing machine started making noise from the drum bearing. Found a Youtube tutorial for bearing replacement on the exact same model, but 3 years older. Followed it just fine until it was time to split the drum. Then it turned out that in the newer units like mine, some rent-seeking MBA fuckers had decid…
But once it knows it’s pretty certain to become common knowledge almost instantaneously. That’s not possible now. What you learn stays localised to you and may be people 1 degree away from you that’s it.
Re: Helix: A vision-language-action model for generalist humanoid control
#129Re: Helix: A vision-language-action model for generalist humanoid control
#130Until we get robots with really good hands, something I'd love in the interim is a system that uses _me_ as the hands. When it's time to put groceries away, I don't want to have to think about how to organize everything. Just figure out which grocery items I have, what storage I have available, come up with an optimized organization solution, then tell me where to put things, one at a time. I'm cautiously optimistic…
> Manna told employees what to do simply by talking to them. Employees each put on a headset when they punched in. Manna had a voice synthesizer, and with its synthesized voice Manna told everyone exactly what to do through their headsets. Constantly. Manna micro-managed minimum wage employees to create perfect performance.