Live data from Hacker News

Helix: A vision-language-action model for generalist humanoid control

figure.ai

151–160 of 178 posts

Re: Helix: A vision-language-action model for generalist humanoid control

#151
post #144

Earlier quoted context omitted.

There is no such thing as "thing" here. These models are trained such that the given conditions (the visual input and the text prompt) will be continued with a desirable continuation (motor function over time). The only dimension accuracy can apply to is desirability.

You don't think there's any segmentation going on?

Implicitly, maybe. Does that matter if you don't know where?

Re: Helix: A vision-language-action model for generalist humanoid control

#152

Earlier quoted context omitted.

I suppose the next big milestone is Wozniak's Coffee Test: A robot is to enter a random home and figure out how to make coffee with whatever they have.

That could still be decades away.

I don't know... I'm starting to seriously think that is only 5-10 years away.

Re: Helix: A vision-language-action model for generalist humanoid control

#153

Seriously, what's with all of these perceived "high-end" tech companies not doing static content worth a damn. Stop hosting your videos as MP4s on your web-server. Either publish to a CDN or use a platform like YouTube. Your bandwidth cannot handle serving high resolution MP4s. /rant

https://www.youtube.com/watch?v=Z3yQHYNXPws

the official figure yt release vid

Re: Helix: A vision-language-action model for generalist humanoid control

#154

Earlier quoted context omitted.

But once it knows it’s pretty certain to become common knowledge almost instantaneously. That’s not possible now. What you learn stays localised to you and may be people 1 degree away from you that’s it.

How does that work? None of the current AI models can re-train on the fly. How would the inference engine even know if it's a case of new information that needs to be fed back, or just a user that's not following instructions correctly?

This is correct. What I meant to say was that in due course, re-training on the fly will become a norm. Even without on the fly re-training we are looking at a small delta.

Re: Helix: A vision-language-action model for generalist humanoid control

#155
post #51
post #46

I don't know, there has been so many overhyped and faked demos in humanoid robotics space over the last couple years, it is difficult to believe what is clearly a demo release for shareholders. Would love to see some demonstration in a less controlled environment.

Imagine they bring one out to a construction site and they treat the robot as a new rookie guy, go pick up those pipes. That would be an ultimate on the fly test to me.

pick up that can, heh heh heh

Re: Helix: A vision-language-action model for generalist humanoid control

#156
post #140

"Pick up anything: Figure robots equipped with Helix can now pick up virtually any small household object, including thousands of items they have never encountered before, simply by following natural language prompts." If they can do that, why aren't they selling picking systems to Amazon by the tens of thousands?

Most AI startups, like most startups in general, are in the business of selling futures. Much easier to get new seed $$$ once there’s a real hype around your demo. Not saying they aren’t being honest, just pointing out the logic of starting here and working your way up to a huge valuation.

Right. As I point out occasionally, Tesla, as a car company, is overvalued by an order of magnitude.

If they can find suckers who accept that valuation, it's much easier to exit as a billionaire than actually make it work.

Visualize making it work. You build or buy a robot that has enough operating envelope for an Amazon picking station, provide it with an end-effector, and use this claimed general purpose software to control it. Probably just arms; it doesn't need to move around. Movement is handled by Amazon's Kiva-type AGV units.

You set up a test station with a supply of Amazon products and put it to work. It's measured on the basis of picks per minute, failed picks, and mean time before failure. You spend months to years debugging standard robotics problems such as tendon wear, gripper wear, and products being damaged during picking and placing. Once it's working, Amazon buys some units and puts them to work in real distribution centers. More problems are found and solved.

Now you have a unit that replaces one human, and costs maybe $20,000 to make in quantity. Amazon beats you down in price so you get to sell it for maybe $25,000 in quantity. You have to build manufacturing facilities and service depots. Success is Amazon buying 50,000 of them, for total income of $0.25 billion. This probably becomes profitable about five years from now, if it all works.

By which time someone in China, Japan, or Taiwan is doing it cheaper and better.

Re: Helix: A vision-language-action model for generalist humanoid control

#157
post #6

There’s nothing I want more than a robot that does house chores. That’s the real 10x multiplier for humans to do what they do best.

I'd pay $2k for something that folds my laundry reliably. It doesn't need arms or legs, just like my dishwasher doesn't need arms or legs. It just needs to let me dump in a load of clean laundry and output stacks of neatly folded or hung clothing.

There are many services that for ~$3/lb will pick up, wash, dry, fold/hang, and deliver 10lb's of laundry every week for $1,500/yr.

Re: Helix: A vision-language-action model for generalist humanoid control

#158
post #140

Earlier quoted context omitted.

Most AI startups, like most startups in general, are in the business of selling futures. Much easier to get new seed $$$ once there’s a real hype around your demo. Not saying they aren’t being honest, just pointing out the logic of starting here and working your way up to a huge valuation.

Right. As I point out occasionally, Tesla, as a car company, is overvalued by an order of magnitude. If they can find suckers who accept that valuation, it's much easier to exit as a billionaire than actually make it work. Visualize making it work. You build or buy a robot that has enough operating envelope for an Amazon picking station, provide it with an end-effector, and use this claimed general purpose software t…

I don’t even think Elizabeth Holmes actually had this mindset. Most entrepreneurs are actually trying to make a business.

Re: Helix: A vision-language-action model for generalist humanoid control

#159
post #60

At this point, this is enough autonomy to have a set of these guys man a howitzer (read as old stockpiles of weapons we already have). Kind of a scary thought. On one hand, I think the idea of moving real people out of danger in war is a good idea, and as an American i'd want Americans to have an edge... and we can't guarantee our enemies won't take it if we skip it, on the other hand I have a visceral reaction to ma…

This would have made a more interesting demo at least

Re: Helix: A vision-language-action model for generalist humanoid control

#160
This whole thread is just people who didn’t read the technical details or immediately doubt the video’s honesty.

I’m actually fairly impressed with this because it’s one neural net which is the goal, and the two system paradigm is really cool. I don’t know much about robotics but this seems like the right direction.

Post reply on HN