Live data from Hacker News

The Battle over Books3

wired.com

61–70 of 135 posts

Re: The Battle over Books3

#61
post #18

From a preservation angle, how big is Books3? Is it easy enough for mortals to mirror for that unlikely future where it might be possible to self-train reasonably good models from scratch if provisioned with data?

800gb

Where are you getting that figure?

Re: The Battle over Books3

#62
post #28

it would be a shame if it would be possible to train a very useful AI using all the books in the world but corporate greed wouldn’t allow it

It's pretty much the opposite: Big corporates are the ones training the models (and benefiting from it).

Either way, it has to happen, and it will happen. It might as well happen on universally-equitable terms.

Re: The Battle over Books3

#63
post #55

Earlier quoted context omitted.

> And if those creators were instead to try to introduce legislation, the AI companies would risk losing access to content from small creators without the means to sue too. this is the point of a class action suit isn't it? if it turns out training isn't fair use then Microsoft/Google/OpenAI will suddenly have class action suits for billions if not trillions of damages against them ($150,000 damages per willful infri…

Have you never seen the outcome of a class action? They’re all slaps in the wrist, less than speeding tickets, and the action members get like a free hotdog or red bull as compensation if they’re lucky

yes, the lawyers always clean up

the aim would be to make training on copyrighted material legally toxic and render all existing datasets and trained weights unlawful

the damages are simply a bonus

Re: The Battle over Books3

#64
post #57

Earlier quoted context omitted.

> incentive for original works maybe goes down I don't know why you would have that concern. LOTS of people write because they just want to share. Now broaden it out to include speech. LOTS of people talk - it's what we do. Progress in AI will happen because humans like to express themselves. The challenge isn't copyright. It's figuring out how to capture the vast content that just isn't getting captured. Also, this…

> LOTS of people write because they just want to share. You just gave up the "for a living" group, who arguably produce overall better content (of course there are exceptions), and focused on hobbyists. I'd call that a self-defeat.

Most all humans communicate for a living. I don't see your point.

Re: The Battle over Books3

#65
post #57

Earlier quoted context omitted.

> LOTS of people write because they just want to share. You just gave up the "for a living" group, who arguably produce overall better content (of course there are exceptions), and focused on hobbyists. I'd call that a self-defeat.

Most all humans communicate for a living. I don't see your point.

But not all and not most produce intellectual content for a living. And those who do you seem to be ok with ditching because if I read you correctly it's a small loss.

Re: The Battle over Books3

#66
post #28

it would be a shame if it would be possible to train a very useful AI using all the books in the world but corporate greed wouldn’t allow it

It's pretty much the opposite: Big corporates are the ones training the models (and benefiting from it).

Other big corporation hold a large portion of copyrights trained on as well. I tend to side with everyone being able to do it

Re: The Battle over Books3

#67
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

SPAs being difficult to crawl is a feature, after all! :)

Re: The Battle over Books3

#68
post #59
post #56

Earlier quoted context omitted.

Texts were copied even before the printing press. See https://en.m.wikipedia.org/wiki/History_of_copyright for more information.

They were, usually to preserve the work, not distribute it on a large scale because there was no large scale. What's your argument ?

I'd guess the point is that prior to the invention of the printing press, an "author" would be financed through a patron. That patron could monetize their initial investment by either hiring people to copy a work by hand or allowing access to the book.

Once the printing press arrived the patron / publisher needed another way to monetize their initial investment as modes of reproduction became more easy, copyright became more restrictive.

Depending on your side of the copyright argument, it either allows whomever is making the initial investment (publisher, author) to be a generous patron of human creative progress ... or ... it allows the to maintain a monopoly on knowledge and be able to profit off it.

Re: The Battle over Books3

#69
post #49
post #34

What worries me most is that this is likely to only increase the gap between large corporate creator and small independent creators. Large AI companies probably are not too worried about making deals with corporate content creators. Having access to content from a trigger-happy creator is only going to increase their advantage over competitors, after all. And if those creators were instead to try to introduce legisla…

[flagged]

>> Copyright was intended to promote and protect human creativity

> This was never the case. Copyright was always a method to enforce private property as a core concept as it's integral to capitalism functioning.

Yeah, copyright was never intended to protect creativity, it was just intended to allow people to make a living by being creative.

Somehow something about that doesn't seem to quite add up.

> Just like police were never about promoting safety and wellbeing.

Please do a bit of research on what happens when laws stop being enforced.

Also: https://en.wikipedia.org/wiki/History_of_the_Metropolitan_Po...

Re: The Battle over Books3

#70
post #59

Earlier quoted context omitted.

They were, usually to preserve the work, not distribute it on a large scale because there was no large scale. What's your argument ?

I'd guess the point is that prior to the invention of the printing press, an "author" would be financed through a patron. That patron could monetize their initial investment by either hiring people to copy a work by hand or allowing access to the book. Once the printing press arrived the patron / publisher needed another way to monetize their initial investment as modes of reproduction became more easy, copyright bec…

> * or ... it allows the to maintain a monopoly on knowledge and be able to profit off it.*

The length of modern copyright terms is absurd and harmful.

The USA started with 14 (plus optionally another 14) which was better.

Last I checked - quite a few year ago - I think there were academic papers calculating that an "optimal" copyright term is probably around 10-14 years.

Post reply on HN