I lean to the side of AI progress over copyright, though I fear AI becoming stupid when incentive for original works maybe goes down, so I think society needs to figure out some social contract at least, like subsidies to keep new content flowing. For instance blogs will stop posting if chatbots in search grab all the answers directly from their sites bypassing all ad revenue.
The Battle over Books3
21–30 of 135 posts
Re: The Battle over Books3
#22I lean to the side of AI progress over copyright, though I fear AI becoming stupid when incentive for original works maybe goes down, so I think society needs to figure out some social contract at least, like subsidies to keep new content flowing. For instance blogs will stop posting if chatbots in search grab all the answers directly from their sites bypassing all ad revenue.
> incentive for original works maybe goes down I don't know why you would have that concern. LOTS of people write because they just want to share. Now broaden it out to include speech. LOTS of people talk - it's what we do. Progress in AI will happen because humans like to express themselves. The challenge isn't copyright. It's figuring out how to capture the vast content that just isn't getting captured. Also, this…
The blogging scene from the early millennium is now a shadow of its former self, and one of the most often stated reasons for abandoning blogging is “my site just wasn’t getting many views any more”. In a world where AI generated content abounds, there will be even fewer eyeballs on whatever one shares and therefore less feeling of reward for sharing. Moreover, the people still blogging are often loading their content with referral links, because in an economy full of glamorous influencers, even ordinary people are tempted to seek some financial reward for sharing beyond the mere pleasure of it. Less eyeballs due to AI competition means fewer people clicking those referral links.
Re: The Battle over Books3
#23Earlier quoted context omitted.
Laws. And the lawsuits that arise when such theft is discovered.
Easy to work around. Contract someone outside of the jurisdiction to provide a dataset, then it's up to them to deliver it to you. You "weren't aware" of the data source until people outside the organization starts shouting or the police starts asking questions, but the model "doesn't contain any of the data" so you continue shipping the model/product.
Re: The Battle over Books3
#24What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…
Is all the data on the internet from 2010 to 2025 that much more valuable than all the data on the internet from 2010 to 2022? Better data yes, but more data up to the present day? Do we really need that to keep improving AI? GPT is already capable of incredible generalized language understanding and I'd wager we've long since hit diminishing returns from raw internet data. RLHF, fine tuning, and better (and more dat…
If your data or models don't account for those then it can make mistakes. For example, if a model is only trained on modern sources (and does not know about Early Modern English 2nd person pronouns "thy"/"thine"/etc.) then it can easily get confused when determining parts of speech, which then affects other down-stream processing.
Re: The Battle over Books3
#25I wonder how this affects things like the Oxford English Dictionary, wiktionary, and other corpus-based linguistics, which rely on sample sentences of the word usage in order to determine the context. Because language is an evolving thing, it is almost certain that they have referenced sentences from copyrighted sources. E.g. I'm willing to bet that they have the sentence where Cory Doctorow introduces the term "ensh…
It shouldn't affect corpus linguistics but the other way around, because all these things have long ago (even before computers) been legally contested by authors and publishers w.r.t. what can be done to text by dictionary makers and corpora managers without the authors' permission, so unless new laws get passed, the current precedents establishing what's permissible for corpus linguistics would still be valid and al…
Re: The Battle over Books3
#26Earlier quoted context omitted.
Laws. And the lawsuits that arise when such theft is discovered.
Easy to work around. Contract someone outside of the jurisdiction to provide a dataset, then it's up to them to deliver it to you. You "weren't aware" of the data source until people outside the organization starts shouting or the police starts asking questions, but the model "doesn't contain any of the data" so you continue shipping the model/product.
Re: The Battle over Books3
#27Look, I get that unethical corporations using this pirated training data for their artist-usurpation machines is bad, right? But EVERYONE being able to dismiss the rights and wishes of current artists while they work to create artist-usurpation machines of their own? That's not any better! You don't need to "democratize access" to that!
Re: The Battle over Books3
#28Re: The Battle over Books3
#29From a preservation angle, how big is Books3? Is it easy enough for mortals to mirror for that unlikely future where it might be possible to self-train reasonably good models from scratch if provisioned with data?
Re: The Battle over Books3
#30it would be a shame if it would be possible to train a very useful AI using all the books in the world but corporate greed wouldn’t allow it