2. Is it possible to estimate how much of copyrighted material has been used?
Ask HN: Estimation of copyright material used by LLM
1–9 of 9 posts
Re: Ask HN: Estimation of copyright material used by LLM
#22. This is harder as a lot of them don't disclose training sets.
Re: Ask HN: Estimation of copyright material used by LLM
#3There's no easy answer there, hence New York Times v. OpenAI.
Re: Ask HN: Estimation of copyright material used by LLM
#4I think what you're looking for is not "copyrighted material" but material that's both 1) used without permission and 2) outside the scope of fair use. There's no easy answer there, hence New York Times v. OpenAI .
I think sticking a straw in Zlib or AA or LibGen or whatever it is, and drinking until it makes gurgling slurping noises as it hoovers up the dregs at the bottom of the barrel, is far, far removed from “fair use”.
Re: Ask HN: Estimation of copyright material used by LLM
#5Re: Ask HN: Estimation of copyright material used by LLM
#6Re: Ask HN: Estimation of copyright material used by LLM
#7Re: Ask HN: Estimation of copyright material used by LLM
#8For example, most popular textbooks have at least several pirate copies uploaded to the web. Some of them are even in plain sight and Googleable.
Re: Ask HN: Estimation of copyright material used by LLM
#9pretty much everything newer than ~70 years old on the internet is copyrighted, because copywright occurs automatically when you create something (in the US at least). So the answer to #1 is yes.