Earlier quoted context omitted.
> I am yet to see any reasoned argument for why it is far more difficult and will take far longer. For language models specifically, they are trained on data and have historically been improved by increasing the size of the model (by number of parameters) and by the amount and/or quality of training data. We are basically out of new, non-synthetic text to train models on and it’s extremely hard work to come up with n…
>We are basically out of new, non-synthetic text to train models this is not even remotely true. There is an astronomical amount of data siloed by publishers, professional journals etc. that is yet to be tapped. OpenAI is making inroads by making deals with these content owners for access to all that juicy data.
You seem to think these models haven't already been trained on pirated versions of this content, for some reason.