Earlier quoted context omitted.
Babies can outperform any CV model with less data than that.
I strongly doubt that. I think babies have much richer input than any CV system up to date. In my opinion movement is crucial for understanding image.
A human baby learns from uncleaned raw data using far less energy with better generalization than a computer and fuses large amount of data without suffering from dimensionality curses.
I think it is safe to say that human babies are still ahead. for now.