Earlier quoted context omitted.
There is no such thing as "thing" here. These models are trained such that the given conditions (the visual input and the text prompt) will be continued with a desirable continuation (motor function over time). The only dimension accuracy can apply to is desirability.
You don't think there's any segmentation going on?
Helix: A vision-language-action model for generalist humanoid control
151–160 of 178 posts
Re: Helix: A vision-language-action model for generalist humanoid control
#152Earlier quoted context omitted.
I suppose the next big milestone is Wozniak's Coffee Test: A robot is to enter a random home and figure out how to make coffee with whatever they have.
That could still be decades away.
Re: Helix: A vision-language-action model for generalist humanoid control
#153Seriously, what's with all of these perceived "high-end" tech companies not doing static content worth a damn. Stop hosting your videos as MP4s on your web-server. Either publish to a CDN or use a platform like YouTube. Your bandwidth cannot handle serving high resolution MP4s. /rant
the official figure yt release vid
Re: Helix: A vision-language-action model for generalist humanoid control
#154Earlier quoted context omitted.
But once it knows it’s pretty certain to become common knowledge almost instantaneously. That’s not possible now. What you learn stays localised to you and may be people 1 degree away from you that’s it.
How does that work? None of the current AI models can re-train on the fly. How would the inference engine even know if it's a case of new information that needs to be fed back, or just a user that's not following instructions correctly?
Re: Helix: A vision-language-action model for generalist humanoid control
#155I don't know, there has been so many overhyped and faked demos in humanoid robotics space over the last couple years, it is difficult to believe what is clearly a demo release for shareholders. Would love to see some demonstration in a less controlled environment.
Imagine they bring one out to a construction site and they treat the robot as a new rookie guy, go pick up those pipes. That would be an ultimate on the fly test to me.
Re: Helix: A vision-language-action model for generalist humanoid control
#156"Pick up anything: Figure robots equipped with Helix can now pick up virtually any small household object, including thousands of items they have never encountered before, simply by following natural language prompts." If they can do that, why aren't they selling picking systems to Amazon by the tens of thousands?
Most AI startups, like most startups in general, are in the business of selling futures. Much easier to get new seed $$$ once there’s a real hype around your demo. Not saying they aren’t being honest, just pointing out the logic of starting here and working your way up to a huge valuation.
If they can find suckers who accept that valuation, it's much easier to exit as a billionaire than actually make it work.
Visualize making it work. You build or buy a robot that has enough operating envelope for an Amazon picking station, provide it with an end-effector, and use this claimed general purpose software to control it. Probably just arms; it doesn't need to move around. Movement is handled by Amazon's Kiva-type AGV units.
You set up a test station with a supply of Amazon products and put it to work. It's measured on the basis of picks per minute, failed picks, and mean time before failure. You spend months to years debugging standard robotics problems such as tendon wear, gripper wear, and products being damaged during picking and placing. Once it's working, Amazon buys some units and puts them to work in real distribution centers. More problems are found and solved.
Now you have a unit that replaces one human, and costs maybe $20,000 to make in quantity. Amazon beats you down in price so you get to sell it for maybe $25,000 in quantity. You have to build manufacturing facilities and service depots. Success is Amazon buying 50,000 of them, for total income of $0.25 billion. This probably becomes profitable about five years from now, if it all works.
By which time someone in China, Japan, or Taiwan is doing it cheaper and better.
Re: Helix: A vision-language-action model for generalist humanoid control
#157There’s nothing I want more than a robot that does house chores. That’s the real 10x multiplier for humans to do what they do best.
I'd pay $2k for something that folds my laundry reliably. It doesn't need arms or legs, just like my dishwasher doesn't need arms or legs. It just needs to let me dump in a load of clean laundry and output stacks of neatly folded or hung clothing.
Re: Helix: A vision-language-action model for generalist humanoid control
#158Earlier quoted context omitted.
Most AI startups, like most startups in general, are in the business of selling futures. Much easier to get new seed $$$ once there’s a real hype around your demo. Not saying they aren’t being honest, just pointing out the logic of starting here and working your way up to a huge valuation.
Right. As I point out occasionally, Tesla, as a car company, is overvalued by an order of magnitude. If they can find suckers who accept that valuation, it's much easier to exit as a billionaire than actually make it work. Visualize making it work. You build or buy a robot that has enough operating envelope for an Amazon picking station, provide it with an end-effector, and use this claimed general purpose software t…
Re: Helix: A vision-language-action model for generalist humanoid control
#159At this point, this is enough autonomy to have a set of these guys man a howitzer (read as old stockpiles of weapons we already have). Kind of a scary thought. On one hand, I think the idea of moving real people out of danger in war is a good idea, and as an American i'd want Americans to have an edge... and we can't guarantee our enemies won't take it if we skip it, on the other hand I have a visceral reaction to ma…
Re: Helix: A vision-language-action model for generalist humanoid control
#160I’m actually fairly impressed with this because it’s one neural net which is the goal, and the two system paradigm is really cool. I don’t know much about robotics but this seems like the right direction.