Claude/OpenAI etc have taken and continue to take literally all data from the internet, printed books, image, audio, and video humans have created in all of existence without permission to train their models. But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
If you put text ‘on the internet’ you do indeed actively permit others to get it. Why lie? Were you burning down archive.org in years past? Then why lie? It is moreover established that training weights on basically anything is legitimate use. Why repeat lie after lie like this? I don’t like LLM mania either but after reading the ten millionth mind-numbing insult to HN intelligence like this I have to think my mother…
But are you ignoring the literal scanning (and 'burning down' of books) that Anthropic has been found guilty of? Or the torrenting of pirated content en-masse by Meta that there is an active lawsuit over to name just 2 recent examples?
Look at the image and audio/video models especially - they can reproduce everything from Mickey Mouse (the copyrighted one) to making entire Seinfeld episodes with the real cast (both visual likeness and even the actor's voices).