Yes the problem is very hard. Mainly because high DOF generalization is very difficult. We have self driving cars because what are the control inputs? Pedal, brake, steering wheel. This already took many many years. Now for a humanoid robot: An action space that is metaphorically Hilbert. (Physically, yes, obviously) Also, IMO, LLM's can aid the development of robots, but do little beyond a planning, human control in…
> So we'll need neural methods which can be learned Data is a problem. LLMs had the advantage of the whole internet to train on. Robots don’t have that corpus of information. And real time learning seems to be something that everyone in AI is studiously ignoring.
Also there’s imitating humans, via a suitable mapping from the human sensor, control and configuration space to the robot’s. Some groups have gathered video and other data from humans doing tasks, for example with a VR headset.