I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…
How do we know what content LLMs were fed? Isn't that a highly guarded secret?
Won't the quality of the content be paramount to the quality of the generated output or does it not work that way?