Earlier quoted context omitted.
No, the above is absolutely exactly true. Robots folding clothes and tying shoelaces etc is nothing but a tech demo at this point. It's difficult to grok this because if you watch a human folding a t-shirt, you can reliably predict that the same human will fold a different t-shirt just as well, and in fact be perfectly capable of folding a wide variety of other clothes items as well. Not so for robots. With robots, w…
You make some very good points about the ACT-2 if it’s nine garment types that’s still very acceptable but the point you bought up about them training on exactly those pieces of garments is a possibility.
In the left half of the video you can see that the robot is (trying to) exactly match the folds of the human in the right half and it's doing so while folding the exact same garments on the exact same surface.
The right half of the video is not a training demonstration, I don't think, since the robot is trained by teleoperation AFAICT (it needs to because it must use its head-mounted camera to control its movements) but that just underlines the degree to which their training regime is exactly copying the movements of a trainer, on the same garment, in the same environment.
This is a limitation of the training approach, by RL. With RL you learn a mapping between sets of pixels (as in the video that comes in through the robot's camera) and robot actions (as in actuator commands). What that means is that once a policy is trained and the robot is deployed, if the input pixels are significantly different than the input pixels at training the robot doesn't have a policy that matches the input pixels and so it can't find the right actions to take. So they have to keep the training and deployment garments and even the environments the same, or as similar as possible.
You can see some more evidence of this in the video right under the paragraph with the title "Hill-Climbing Reliability Through Post-Training". The robot at the front of the video, with the bright red trim, is shown trying to fold a grey t-shirt with white flower decorations and a frilly hem (how adorable before the robot has completed the fold. If it could complete it, you can rest assured that the video would be showing off the entire folding sequence as it does for the robot with the green cap on the other side of the bed.
That paragraph is making a claim about a "post-training" regime that's supposed to improve generalisation but it leaves more details to a "separate technical post". So I can't tell what it's supposed to be doing, but I don't think it's working.
When I watch videos like that I always remind myself that a) I'm watching a tech demo created to attract investment and b) I've watched way too many of those, going all the way back to the Boston Dynamic videos of Robot Dog or of Atlas doing backflips and yet the state of the art hasn't really budged since. Such videos make it easy to overestimate the state of the art in autonomous robotics and in fact are meant do precisely that: play up robots' true capabilities. It's just impossible to say anything about a robot's general capabilities by watching a few minutes or even a few hours of video. OtoH if you know what to look for you can tell everyone is basically stuck at the same level and trying the same things to escape it. The truth is robotic autonomy is several major breakthroughs away and nobody has any idea how to get there. So we'll be seeing many more of those tech demo videos in the years to come.