Live data from Hacker News

The Battle over Books3

wired.com

1–10 of 135 posts

Re: The Battle over Books3

#2
What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this?

The problem with my question is the following:

Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to other sites, which are unlikely to block GPTBot themselves. How will DigitalOcean (or any other company) address this?

It looks to me like it's Catch 22 in every direction you look, and unless you're someone like The New York Times who can afford to outright protect the data with licensing...

It's just something thats been on my mind lately but I don't understand the finer details of it.

Re: The Battle over Books3

#3
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

My guess is the big players hope is to steal an enough content and then build a self training LLM based off synthetic content (rehashed original works) before the theft part matters. Not sure how close they are to achieving but this seems to be a common SV gamble.

Steal or do something shady, raise enough money / power so by the time your noticed, you have the money to win in the courts.

You already see the propaganda about “none of this should matter because the cure for cancer is on the way courtesy of AI.” I mean maybe it it but it smells fishy to me.

Re: The Battle over Books3

#4
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

Laws. And the lawsuits that arise when such theft is discovered.

Re: The Battle over Books3

#5
Does anyone have more content about what makes Books3 so special relative to Bibliotik? Was it processed somehow, or just compiled into a single file?

I feel like I’m missing something from this and all of the other articles about Books3. It sounds like he downloaded all of the books from a book piracy site then rehosted them with the “Books3” name. Surely there must be more to the story? Or is the story simply that a professor hosted pirated content under his own name under the guise of AI training?

This is the kind of effort that could have been done anonymously, just as all of the pirated books had already been uploaded and hosted anonymously. I’m not sure why he expected any different outcome by re-pirating everything under his own name.

The journalists seem to be loving it, though. All of the tech journals have an article about this guy.

Re: The Battle over Books3

#6
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

It's going to be a walled garden dystopia, with everyone and everything asking to sign-up, subscribe or whatever to monetize content.

Scraping & siphoning content, ad blockers, both sides have been in arms race for years, AI and LLMs will just be the last straw, before almost anything of value is behind a pay/subscription wall.

Re: The Battle over Books3

#7
I lean to the side of AI progress over copyright, though I fear AI becoming stupid when incentive for original works maybe goes down, so I think society needs to figure out some social contract at least, like subsidies to keep new content flowing.

For instance blogs will stop posting if chatbots in search grab all the answers directly from their sites bypassing all ad revenue.

Re: The Battle over Books3

#8
I wonder how this affects things like the Oxford English Dictionary, wiktionary, and other corpus-based linguistics, which rely on sample sentences of the word usage in order to determine the context.

Because language is an evolving thing, it is almost certain that they have referenced sentences from copyrighted sources. E.g. I'm willing to bet that they have the sentence where Cory Doctorow introduces the term "enshittification". (The OALD 7ed in it's foreword even states "Corpus analysis now makes it possible to draw authentic examples from a vast range of attested contemporary usage. A concordance will display hundreds or thousands of them to choose from.")

I suspect that the inclusion of a few sentences -- especially those that introduce a new word or usage of a word -- are fair use, but the inclusion of the entire texts is not.

This then brings up an interesting point where the computer scientists/linguists developing tools like WordNet or other NLP databases would be at an advantage to those that take the approach of throwing a lot of data into a neural network and hoping for the best. Yes, it is a lot more work/effort to develop those NLP databases, but in the end they may end up being more robust, especially around the question of copyright.

Re: The Battle over Books3

#9
post #3
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

My guess is the big players hope is to steal an enough content and then build a self training LLM based off synthetic content (rehashed original works) before the theft part matters. Not sure how close they are to achieving but this seems to be a common SV gamble. Steal or do something shady, raise enough money / power so by the time your noticed, you have the money to win in the courts. You already see the propagand…

An LLM can be trained to find relevant knowledge online. It doesn’t have to be trained on the entirety of all existing knowledge.

Re: The Battle over Books3

#10
post #4
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

Laws. And the lawsuits that arise when such theft is discovered.

Easy to work around. Contract someone outside of the jurisdiction to provide a dataset, then it's up to them to deliver it to you. You "weren't aware" of the data source until people outside the organization starts shouting or the police starts asking questions, but the model "doesn't contain any of the data" so you continue shipping the model/product.
Post reply on HN