Live data from Hacker News

Helix: A vision-language-action model for generalist humanoid control

figure.ai

171–178 of 178 posts

Re: Helix: A vision-language-action model for generalist humanoid control

#171
post #170
post #67

I'm always wondering at the safety measures on these things. How much force is in those motors? This is basically safety-critical stuff but with LLMs. Hallucinating wrong answers in text is bad, hallucinating that your chest is a drawer to pull open is very bad.

That's actually more of a solved problem. Robot arms that can track the force they're applying and where to avoid injuring humans have been kicking around for 10-15 years. It let them go out of the mega safety cells into the same space as people and even do things like letting the operator pose the robot to teach it positions instead of having to do it in a computer program or with a remote control. The term I see a…

That's fine for wheeled robots or robots bolted to the floor but for legged robots, especially bipeds, the hard question is how to prevent them from falling over on things. These don't look heavy enough to be too dangerous for a standing adult but you've still got pets/children to worry about.

Re: Helix: A vision-language-action model for generalist humanoid control

#172

Earlier quoted context omitted.

That is how I would design it. It is common in safety critical PLC systems to have 1 or more separate safety PLCs that try to prevent bad things from happening.

Although in a SIL safety system the dangerous events are identified and extremely thoroughly characterized as part of system design. There cannot be a safety system of this type for a generalist platform like a humanoid robot. It's possibility space is just too high. I think the safety governor in this case would have to be a neural network that is at least as complex as the robots network, if not more so. Which begs…

Limiting max force applied CAN be can be characterized for this robot.

Re: Helix: A vision-language-action model for generalist humanoid control

#173

Earlier quoted context omitted.

You can have dedicated controllers for the motors that limit their max torque.

That's not enough. When a robot link is in motion and hits an object, the change in momentum creates an impulse over the duration of deceleration. The faster the robot moves, the faster it has to decelerate, the higher the instantaneous braking force at the impact point.

Then limit max velocity.

Re: Helix: A vision-language-action model for generalist humanoid control

#174
post #70
post #6

There’s nothing I want more than a robot that does house chores. That’s the real 10x multiplier for humans to do what they do best.

To me this is such a weird wish. Why would you not want to care for your home and the people living there? Why would you want to have a slave taking these activities from you? I'd rather have less waged labour and more time for chores with the family.

Because it's boring and tedious. And no, a robot is not a slave.

Re: Helix: A vision-language-action model for generalist humanoid control

#175
post #60

At this point, this is enough autonomy to have a set of these guys man a howitzer (read as old stockpiles of weapons we already have). Kind of a scary thought. On one hand, I think the idea of moving real people out of danger in war is a good idea, and as an American i'd want Americans to have an edge... and we can't guarantee our enemies won't take it if we skip it, on the other hand I have a visceral reaction to ma…

We already have an ongoing major war (in Ukraine) where both sides are using autonomous AI-driven drones that kill people, at scale, with escalating tit-for-tat advances in lethality. This conversation is being had alright.

Re: Helix: A vision-language-action model for generalist humanoid control

#176
post #87
post #68

Earlier quoted context omitted.

Not a big deal on the battlefield.

I'd say a very big deal when munitions and targeting are involved

Why do you think that?

Battle fields are violent and exhausting, people get the shakes, make mistakes, hurt themselves and each other all the time. Munitions are generally designed to not explode from a bump in the road or getting dropped or squeezed, and targeting systems commonly support automatic tracking or similar, for this specific reason.

Re: Helix: A vision-language-action model for generalist humanoid control

#177
post #116

The demo is quite interesting but I am mostly intrigued by the claim that it is running totally local to each robot. It seems to use some agentic decision making but the article doesn't touch on that. What possible combo of model types are they stringing together? Or is this something novel? The article mentions that the system in each robot uses two ai models. S2 is built on a 7B-parameter open-source, open-weight V…

I'm very skeptical. I'm quite familiar with VLAs and this seems like an unbelievable leap forward based on their claims.

Re: Helix: A vision-language-action model for generalist humanoid control

#178

Earlier quoted context omitted.

Black Mirror has the perfect episode for this scenario already.

Except those things would be very easily defeated by a 12 gauge shotgun or a AR-15

You are severely underestimating the speed of drones and their formation capabilities, as well as overestimating your aim under immense danger.
Post reply on HN