>Furthermore, the fact that LLMs seem to need such a stupendous amount of data to get such mediocre reasoning indicates that they simply are not generalizing. If these models can’t get anywhere close to human level performance with the data a human would see in 20,000 years, we should entertain the possibility that 2,000,000,000 years worth of data will be also be insufficient. There’s no amount of jet fuel you can a…
We haven’t had ML models this large before. There’s innovation in architecture but we often come back to the bitter lesson: more data.
We’re likely going to see experimentation with language models to learn from few examples. Fine tuning pretrained LLMs shows they have quite a remarkable ability to learn from few examples.
Liquid AI has a new learning architecture for dynamic learning and much smaller models.
Some people seem mad about the bitter lesson, they want their model based on human features to work better when so far usually more data wins.
I think the next evolution here is in increasing the quality of the training data and giving it more structure. I suspect the right setup can seed emergent capabilities.