Live data from Hacker News

Publishers want billions, not millions, from AI

semafor.com

21–30 of 78 posts

Re: Publishers want billions, not millions, from AI

#21

So this was the outcome I guessed would need to happen, but I didn't think publishers would be the ones to push for it, but I guess they are. The thing is if this upsets you on some gut level, this is just a sign for how valuable information to train on is. May be billions is too much, but hundreds of millions? If your immediate thought is "this must be stopped, this isn't fair" it is a signal of how reliant AI at al…

Exactly right. The leaked Google AI "no moat" memo applies to the technology itself, but not the data.

Without formal agreements with content owners, OpenAI is just another scraper. One can easily imagine that content cooperatives may ultimately own the GenAI space.

Re: Publishers want billions, not millions, from AI

#23
post #15

I am reminded somewhat of the ride & delivery apps here. While they did use tech to enable some cool things like demand pricing and efficient route planning, a big part of their innovation really came from using that tech to shovel most of the risk and cost of providing the service onto independent contractors. LLMs are a genuinely exciting technology, but I am worried that a part of what they enable will turn out to…

> If you train your AI on a bunch of someone elses work, you can make it produce something very simiilar As a blogger of 10+ years this has been very interesting to realize. I can use ChatGPT to write really good first drafts because “in the style of Swizec Teller” works as a prompt. The results are way better than the usual generic corporate drone style of output. BUT! And this is an important but. The insights it p…

> Every sharp insight gets dulled by “AI protections”, every hot take is two-sided to death, etc.

This is a OpenAI thing.

Run a "uncensored" llama finetune, like airoboros 70b. They are as hot and spicy as you want them to be.

EDIT: Actually someone is hosting 65b here, so you can test it and see what I mean: https://lite.koboldai.net/

Note that you have to get the instruct prompting syntax right.

Re: Publishers want billions, not millions, from AI

#24
post #17

Big media companies go fuck yourself. That nightmare scenario, for Levin, would turn a Food & Wine review into a simple text recommendation of a bottle of Malbec. If there is just one wine review on the web which recommends Malbec, AI will not start recommending it. If there are many such reviews, then yes, AI will tell you "Many reviews recommend Malbec.". Just like a human can tell you about what they learned, our…

AI is not my “brother” nor is it a human. This is a predictive model controlled by a for profit business and yea, they should be paying out if they are training on data owned by someone.

It is not a human but it is a fellow intelligent system.

Mankind had this discussion before. When Darwin published his theory about evolution. He faced a lot of hatred because it made humans less special.

Now we go through the same dance again. This time, intelligence in silicon makes humans less special.

Re: Publishers want billions, not millions, from AI

#25

Hilarious and desperate. They are in no position to make demands. The current data that has been scraped for AI models is probably good enough to be able to generate synthetic data. Even if it's not, what are publishers going to do about people scraping their content? Absolutely nothing. Trying to force this issue feels like a good way to be disintermediated quicker .

To hell with writers, I for one look forward to our brave new world of fake plastic trees.

Large publishers are not writers. It would probably be better if large publishers died and there were more small publishers or independent writers.

Re: Publishers want billions, not millions, from AI

#26

I am reminded somewhat of the ride & delivery apps here. While they did use tech to enable some cool things like demand pricing and efficient route planning, a big part of their innovation really came from using that tech to shovel most of the risk and cost of providing the service onto independent contractors. LLMs are a genuinely exciting technology, but I am worried that a part of what they enable will turn out to…

Acknowledging these LLMs are not human, but isn’t this kinda what humans do? Take in lots of different examples and produce something similar but distinctly different and not paying royalties or being considered plagiarism.

Yeah, I mean we have this term “inspiration” already for some time. I’ve long thought that when a person hits a dead end creatively what they need is to to go out and have more experiences and be exposed to new things.

Re: Publishers want billions, not millions, from AI

#27
post #5
post #2

Presumably OpenAI and others (for the most part) are checking the licenses for content they use in training? Let’s say for example that they trained on the content of the entire archive of the New York Times… isn’t it safe to say they’d have purchased a license for that content from NYT? Wouldn’t “commercial use” cover this? Seems like the publishers just want a piece of the money because they want a piece of the mon…

OpenAI absolutely does not have licenses for 99%+ of the content used to build their models. They're following the standard tech company model of "negotiate forgiveness rather than ask for permission".

Exactly. When you have enough money, you think laws are for other people, as with Uber or cryptocurrency.

The legal status of copyright in this context is unclear, because the laws were written without imitation machines in mind. Presumably they blew right on by "is this right?" and went for "can we get away with this and pay less in fines than we'd make in profits?" As long as there's a 51% chance of that being the case, they saw it as worth a go.

Re: Publishers want billions, not millions, from AI

#28
post #2

Presumably OpenAI and others (for the most part) are checking the licenses for content they use in training? Let’s say for example that they trained on the content of the entire archive of the New York Times… isn’t it safe to say they’d have purchased a license for that content from NYT? Wouldn’t “commercial use” cover this? Seems like the publishers just want a piece of the money because they want a piece of the mon…

I don't think that's a safe presumption at all. Getty, e.g., claims to have found their watermark in the output of Stable Diffusion.

More broadly, it's not a settled question of what license you need. The article presents a perspective that LLMs "would turn a Food & Wine review into a simple text recommendation of a bottle of Malbec, without attribution." I doubt that anyone has a license to republish the entire archive of the New York Times without attribution.

At the other extreme, I can "train" myself on newspapers I find discarded on the subway with no license at all. I don't need to cite my sources when I state an opinion based on everything I've ever read, seen, or heard. Is that a fair analogy to apply to LLMs?

Re: Publishers want billions, not millions, from AI

#29
Ah, publisher rent seeking. To get independent again. Like they were in the haydays of the cold War or the Iraq war.not.

Why not a different rent seeking model.. For pharma companies? If your product cures a patient, all other pharma companies own you a percentage from everything sold ever after.

Re: Publishers want billions, not millions, from AI

#30
post #5

Earlier quoted context omitted.

OpenAI absolutely does not have licenses for 99%+ of the content used to build their models. They're following the standard tech company model of "negotiate forgiveness rather than ask for permission".

Yep, my understanding is that one of their datasets is basically the ebook dump of z-lib. Honestly, they're likely to get away with training on copyrighted work unless someone can get the LLMs to spit out whole pages of copyrighted material. I don't even think small excerpts would be breaking fair use

This is big enough, though, that we'll likely see new laws made in response, the same way the rise of the internet caused previously ambiguous things to get clear laws around it. So I think they also thought they could force the ambiguity to resolve in their direction with a combination of "facts on the ground" and a big budget for lobbying and litigation.
Post reply on HN