Live data from Hacker News

The Battle over Books3

wired.com

81–90 of 135 posts

Re: The Battle over Books3

#81
post #80

Earlier quoted context omitted.

> Which was not at scale I doubt you are playing here with a good-faith definition of “at scale”. In any event, copying of works for sale in Antiquity was certainly of scale, we know that many literary works spread quickly across the Mediterranean through commercial production. Furthermore, some of the earliest printed books were made in such limited editions that Roman mass production can certainly be compared. The…

> I doubt you are playing here with a good-faith definition of “at scale”. I could say the same thing about you because you make it sound like the appearance of some works in various places on the map is equivalent to their vast abundance. > development of the printing press is generally viewed in contrast to the medieval manuscript era that immediately preceded it If you skip renaissance. > As for going to a show an…

> the appearance of some works in various places on the map is equivalent to their vast abundance.

As I said, historians know that some popular works not only appeared across the map quickly, they were commercially sold in the marketplace such that the educated class was able to purchase their own copies of prominent recent works with, of course, no money going back to the creator. Again, I don’t think your definition of “abundance” is good faith.

With regard to your last point, why should the artists’ desire for compensation outweigh the desire of audiences to consume the media for free, or other artists’ desire to rework prior art for free? This is a moral debate that is quite culturally dependent, and though you want to claim that your views are the right ones, that just won’t fly on a forum as international as HN. Many posters on HN grew up with pirated DVD and cassette/CD stands at the market (some might even still have them where they live), or in their countries Bittorrent or now pirate streaming sites are things used by ordinary people.

Re: The Battle over Books3

#83
is there enough data on the web for a LLM to be agnostic about the source languages, Russian, Chinese, English, German, etc? Where training on Russian and Chinese and English and German etc sources would also incorporate enough information about translation that if the AI learned about some topic only through Chinese sources, it could still recognize/use/apply/express those ideas in English?

Re: The Battle over Books3

#84
post #80

Earlier quoted context omitted.

> I doubt you are playing here with a good-faith definition of “at scale”. I could say the same thing about you because you make it sound like the appearance of some works in various places on the map is equivalent to their vast abundance. > development of the printing press is generally viewed in contrast to the medieval manuscript era that immediately preceded it If you skip renaissance. > As for going to a show an…

> the appearance of some works in various places on the map is equivalent to their vast abundance. As I said, historians know that some popular works not only appeared across the map quickly, they were commercially sold in the marketplace such that the educated class was able to purchase their own copies of prominent recent works with, of course, no money going back to the creator. Again, I don’t think your definitio…

Ok, I'd like to call you on that and provide actual numbers. I found this:

Before the invention of printing, the number of manuscript books in Europe could be counted in thousands. By 1500, after only 50 years of printing, there were more than 9,000,000 books. https://www.britannica.com/topic/publishing/The-medieval-boo...

Which illustrates what the difference in scale I'm talking about.

> why should the artists’ desire for compensation outweigh the desire of audiences to consume the media for free

Because if you disincentivize the artist there are no media to consume. Why should your desire to consume for free deprive me from consuming at all if there is no artist willing to accept such conditions ?

Re: The Battle over Books3

#85
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

Is all the data on the internet from 2010 to 2025 that much more valuable than all the data on the internet from 2010 to 2022? Better data yes, but more data up to the present day? Do we really need that to keep improving AI? GPT is already capable of incredible generalized language understanding and I'd wager we've long since hit diminishing returns from raw internet data. RLHF, fine tuning, and better (and more dat…

People involved in AI have incentive to say 'yes' to everything.

Nobody will make anything value in public, that they didn't want to be released for free anyway.

ChatGPT's vacuum has brought back a desire for privacy and will probably contribute to destroying piracy too.

ChatGPT has destroyed the 'study hard and get reward loop' for collaborative effort on the internet. If you use chatGPT, it absorbs all your question data and gives you nothing in return. You can't commit to random people, as they are expected to leak your IP onto gpt.

Isaac Newton using chatGPT would upload the core of calculus to GPT in research questions, and see no personal benefit for doing so.

There is no greater thief in history of academic work, than electronics.

Re: The Battle over Books3

#86
post #82

“I was poking around, Googling ‘how to download Library Genesis,’” LG literally has this info on its front page. Researchers will research, I suppose.

More to the point, it’s huge. 33 TB is a massive amount, and I was grateful to sidestep it by downloading just the epubs.

Still, there’s always room for books4.

Re: The Battle over Books3

#87
post #83

is there enough data on the web for a LLM to be agnostic about the source languages, Russian, Chinese, English, German, etc? Where training on Russian and Chinese and English and German etc sources would also incorporate enough information about translation that if the AI learned about some topic only through Chinese sources, it could still recognize/use/apply/express those ideas in English?

That’s a fascinating question. I’m not sure. It seems like if it learned about a topic in Chinese, it would be able to express it in English, but I haven’t seen this tested.

Re: The Battle over Books3

#88
post #28

it would be a shame if it would be possible to train a very useful AI using all the books in the world but corporate greed wouldn’t allow it

What if I don't want my book used for that purpose?

It isn't as though the AI companies have even paid for a single copy of the authors' books.

Re: The Battle over Books3

#89
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

I did a random sample of some news sites

across the political spectrum: every single one I checked, other than ft.com blocks GPTBot

Re: The Battle over Books3

#90
post #63

Earlier quoted context omitted.

Have you never seen the outcome of a class action? They’re all slaps in the wrist, less than speeding tickets, and the action members get like a free hotdog or red bull as compensation if they’re lucky

yes, the lawyers always clean up the aim would be to make training on copyrighted material legally toxic and render all existing datasets and trained weights unlawful the damages are simply a bonus

Yeah that's the real goal. You can make it impossible for big corporations to use this technology because it paints a huge target on them while ignoring the de facto state of open source models that are impossible to enforce regulation against. Best of both worlds maybe.
Post reply on HN