Succinctly stated and something that resonates strongly with me. In the last internet revolution (web search), results started high quality because the inputs were high quality - bloggers and others just wanted to document and share knowledge. But over time, many interests (largely commercial) figured out how to game the system with SEO, and quality of search results has decreased as search's incentive structure led…
If the AI was good at detecting it, it wouldn't matter if the AI-generated content sucked, yes?
Even a low probability of detection would help. Let's say our algorithm is 50% likely to detect AI junk. That means that half the junk data won't make it in to model. Even 20% would probably be worthwhile, especially if it also threw away human-generated junk (and let's be realistic here: there is, and always has been, no shortage of terrible and/or wrong human-generated content).
Let's say those crappy filler paragraphs that get stuck between pictures on meme clickbait pages... I suspect most of those people have already been replaced, but that prose was horrible long before the current LLM boom.
I suspect there is a lot of effort being expended right now on ways to ensure the training data (whether human or AI generated) isn't shite.