Earlier quoted context omitted.
I think we hit peak AI improvement velocity sometime mid last year. The reality is all progress was made using a huge backlog of public data. There will never be 20+ years of authentic data dumped on the web again. I've hoped against but suspected that as time goes on LLMs will become increasingly poisoned by the the well of the closed loop. I don't think most companies can resist the allure of more free data as bitt…
> I don't think most companies can resist the allure of more free data as bitter as it may taste. Mercor, Surge, Scale, and other data labelling firms have shown that's not true. Paid data for LLM training is in higher demand than ever for this exact reason: Model creators want to improve their models, and free data no longer cuts it.
Doesn't change my point, I still don't think they can resist pulling from the "free" data. Corps are just too greedy and next quarter focused.