US publishers tell Common Crawl to stop scraping and delete archive
pressgazette.co.uk
US publishers tell Common Crawl to stop scraping and delete archive
1–10 of 11 posts
Re: US publishers tell Common Crawl to stop scraping and delete archive
#2Re: US publishers tell Common Crawl to stop scraping and delete archive
#3It's absurd to say "you can't record this book to a friend or robot".
Nobody seems to actually reproduce the copyrighted materials.
High-dimensional eigendecompositions which underpin AI similarity are some of the most literally derivative materials of texts that you can imagine.
Re: US publishers tell Common Crawl to stop scraping and delete archive
#4Re: US publishers tell Common Crawl to stop scraping and delete archive
#5Re: US publishers tell Common Crawl to stop scraping and delete archive
#6This is shady. Copyrighters absolutely not get to control use of their copyrighted material when people mentally, sonically, or physically reproduce it for personal use. It's absurd to say "you can't record this book to a friend or robot". Nobody seems to actually reproduce the copyrighted materials. High-dimensional eigendecompositions which underpin AI similarity are some of the most literally derivative materials…
(my point being that it would be different if the product CommonCrawl provides were trained models, but this is not the case: its product is unlawful reproductions of copyrighted data for commercial use)
Re: US publishers tell Common Crawl to stop scraping and delete archive
#7Its like they don't understand the problem common crawl solved rather neatly. You think the skid scrapers are bad? Wait till the competent players lose access to CC.
Re: US publishers tell Common Crawl to stop scraping and delete archive
#8The publishers need to rethink their entire take on how the Internet works or any "victory" they earn is going to be extremely Pyrrhic.
Re: US publishers tell Common Crawl to stop scraping and delete archive
#9This is shady. Copyrighters absolutely not get to control use of their copyrighted material when people mentally, sonically, or physically reproduce it for personal use. It's absurd to say "you can't record this book to a friend or robot". Nobody seems to actually reproduce the copyrighted materials. High-dimensional eigendecompositions which underpin AI similarity are some of the most literally derivative materials…
So you record a copy "for a friend" and then you sell lots of those copies as your business. All within your rights! What's mine is yours, my Comrade! (my point being that it would be different if the product CommonCrawl provides were trained models, but this is not the case: its product is unlawful reproductions of copyrighted data for commercial use)
Common Crawl is not a business and is not selling anything.
Re: US publishers tell Common Crawl to stop scraping and delete archive
#10Earlier quoted context omitted.
So you record a copy "for a friend" and then you sell lots of those copies as your business. All within your rights! What's mine is yours, my Comrade! (my point being that it would be different if the product CommonCrawl provides were trained models, but this is not the case: its product is unlawful reproductions of copyrighted data for commercial use)
> then you sell lots of those copies as your business Common Crawl is not a business and is not selling anything.