Earlier quoted context omitted.
> "No human requires the amount of training data we are already giving the models." Well, humans are also trained differently. We interact with other humans in real time and get immediate feedback on their responses. We don't just learn by reading through reams of static information. We talk to people. We get into arguments. And so on. Maybe the ideal way to train an AI is to have it interact with lots of humans, so…
Yeah I'm not totally convinced humans don't have a tremendous amount of training data - interacting with the world for years with constant input from all our senses and parental corrections. I bet if you add up that data it's a lot. But once we are partially trained, training more requires a lot less.
Let's Fermi estimate that.
A 4k video stream is about 50 megabits/second. Let's say that humans have the equivalent of two of those going during waking hours, one for vision and one for everything else. Humans are awake for 18 hours/day, and we'll say a human's training is 'complete' at 25.
Multiply that together, and you end up with 1.8e17 bytes, or 180 petabytes of data.
There's plenty of reason to think that we don't learn effectively from all of this data (there's lots of redundancy, for example), but at the grossest orders of magnitude you seem to be right that at least in theory we have access to a tremendous amount of data.