Live data from Hacker News

Publishers want billions, not millions, from AI

semafor.com

11–20 of 78 posts

Re: Publishers want billions, not millions, from AI

#11

So this was the outcome I guessed would need to happen, but I didn't think publishers would be the ones to push for it, but I guess they are. The thing is if this upsets you on some gut level, this is just a sign for how valuable information to train on is. May be billions is too much, but hundreds of millions? If your immediate thought is "this must be stopped, this isn't fair" it is a signal of how reliant AI at al…

[dead]

Re: Publishers want billions, not millions, from AI

#12
Hilarious and desperate.

They are in no position to make demands. The current data that has been scraped for AI models is probably good enough to be able to generate synthetic data.

Even if it's not, what are publishers going to do about people scraping their content? Absolutely nothing.

Trying to force this issue feels like a good way to be disintermediated quicker.

Re: Publishers want billions, not millions, from AI

#13
post #2

Presumably OpenAI and others (for the most part) are checking the licenses for content they use in training? Let’s say for example that they trained on the content of the entire archive of the New York Times… isn’t it safe to say they’d have purchased a license for that content from NYT? Wouldn’t “commercial use” cover this? Seems like the publishers just want a piece of the money because they want a piece of the mon…

> Presumably OpenAI and others (for the most part) are checking the licenses for content they use in training?

Ha.

There is a difficult to describe culture in the ML community, from the top researchers to casual tinkerers... Everything is bleeding edge and moving at blinding speed. Some devs don't even stop to think about things like licensing or easy reproducibility/packaging, because they are too busy moving onto the next thing.

This thread on how a popular repo was unlicensed and violating other licenses for months is a good example: https://github.com/AUTOMATIC1111/stable-diffusion-webui/issu...

OpenAI comes from this culture, even if they are a more commercial company now.

Re: Publishers want billions, not millions, from AI

#14

Hilarious and desperate. They are in no position to make demands. The current data that has been scraped for AI models is probably good enough to be able to generate synthetic data. Even if it's not, what are publishers going to do about people scraping their content? Absolutely nothing. Trying to force this issue feels like a good way to be disintermediated quicker .

To hell with writers, I for one look forward to our brave new world of fake plastic trees.

Re: Publishers want billions, not millions, from AI

#15

I am reminded somewhat of the ride & delivery apps here. While they did use tech to enable some cool things like demand pricing and efficient route planning, a big part of their innovation really came from using that tech to shovel most of the risk and cost of providing the service onto independent contractors. LLMs are a genuinely exciting technology, but I am worried that a part of what they enable will turn out to…

> If you train your AI on a bunch of someone elses work, you can make it produce something very simiilar

As a blogger of 10+ years this has been very interesting to realize. I can use ChatGPT to write really good first drafts because “in the style of Swizec Teller” works as a prompt. The results are way better than the usual generic corporate drone style of output.

BUT! And this is an important but. The insights it produces remain at that generic corporate drivel level. Even with a full list of bullet points of cool insights, an LLM takes those and averages them out into a big pile of meh. Every sharp insight gets dulled by “AI protections”, every hot take is two-sided to death, etc.

Despite superficially following my style, the AI seems unable to catch your attention. Everything it writes is a little boring, a little too long, a dash tiring to read. The spark is missing.

Also my style has evolved since 2021 and that’s missing.

Re: Publishers want billions, not millions, from AI

#16
post #6

I am out of sympathy for big publishers. They have been greedy and predatory for decades, at the expense of their own content.

Yes, I agree, they could have invested in the various AI companies, instead the kept bleeding their content creators of all they have. Well the boat sailed, see you later big publishers.

> they could have invested in the various AI companies

Not really what I meant.

> instead the kept bleeding their content creators of all they have

But yes, that. And their customers too.

Re: Publishers want billions, not millions, from AI

#17
Big media companies go fuck yourself.

    That nightmare scenario, for Levin,
    would turn a Food & Wine review into
    a simple text recommendation of a
    bottle of Malbec.
If there is just one wine review on the web which recommends Malbec, AI will not start recommending it. If there are many such reviews, then yes, AI will tell you "Many reviews recommend Malbec.". Just like a human can tell you about what they learned, our artificial brothers have the same right to learn.

Nobody should be allowed to bend the law and extort money from society for writing wine reviews.

Re: Publishers want billions, not millions, from AI

#18
post #17

Big media companies go fuck yourself. That nightmare scenario, for Levin, would turn a Food & Wine review into a simple text recommendation of a bottle of Malbec. If there is just one wine review on the web which recommends Malbec, AI will not start recommending it. If there are many such reviews, then yes, AI will tell you "Many reviews recommend Malbec.". Just like a human can tell you about what they learned, our…

AI is not my “brother” nor is it a human. This is a predictive model controlled by a for profit business and yea, they should be paying out if they are training on data owned by someone.

Re: Publishers want billions, not millions, from AI

#19

I am reminded somewhat of the ride & delivery apps here. While they did use tech to enable some cool things like demand pricing and efficient route planning, a big part of their innovation really came from using that tech to shovel most of the risk and cost of providing the service onto independent contractors. LLMs are a genuinely exciting technology, but I am worried that a part of what they enable will turn out to…

Acknowledging these LLMs are not human, but isn’t this kinda what humans do? Take in lots of different examples and produce something similar but distinctly different and not paying royalties or being considered plagiarism.

Re: Publishers want billions, not millions, from AI

#20
post #5
post #2

Presumably OpenAI and others (for the most part) are checking the licenses for content they use in training? Let’s say for example that they trained on the content of the entire archive of the New York Times… isn’t it safe to say they’d have purchased a license for that content from NYT? Wouldn’t “commercial use” cover this? Seems like the publishers just want a piece of the money because they want a piece of the mon…

OpenAI absolutely does not have licenses for 99%+ of the content used to build their models. They're following the standard tech company model of "negotiate forgiveness rather than ask for permission".

Yep, my understanding is that one of their datasets is basically the ebook dump of z-lib. Honestly, they're likely to get away with training on copyrighted work unless someone can get the LLMs to spit out whole pages of copyrighted material. I don't even think small excerpts would be breaking fair use
Post reply on HN