Data requirements are overstated - you can train on longer and longer sequences and I am pretty sure most organizations are still using the “show the model the data only once“ approach which is just wasteful. Compute challenges are more real, but we are seeing for the first time huge amounts of global capital being allocated to solve specifically these problems, so I am curious what fruit that will bear in a few year…
> using the “show the model the data only once“ approach which is just wasteful. According to the InstructGPT paper, that is not the case, showing the data multiple times results in overfitting.
2. They still saw performance improvements which is why they did train on the data multiple times, you can see in the paper.
3. there was a recent paper demonstrating that reusing data still saw continued improvements in perplexity, i am on my ipad so cannot find it now