Publishers want billions, not millions, from AI
1–10 of 78 posts
Re: Publishers want billions, not millions, from AI
#2Let’s say for example that they trained on the content of the entire archive of the New York Times… isn’t it safe to say they’d have purchased a license for that content from NYT?
Wouldn’t “commercial use” cover this?
Seems like the publishers just want a piece of the money because they want a piece of the money.
Re: Publishers want billions, not millions, from AI
#3Re: Publishers want billions, not millions, from AI
#4The thing is if this upsets you on some gut level, this is just a sign for how valuable information to train on is. May be billions is too much, but hundreds of millions? If your immediate thought is "this must be stopped, this isn't fair" it is a signal of how reliant AI at all is on the data it trains. An untrained nn is useless, it's only useful if it has information to train on. It's not like value creation is zero sum but it certainly doesn't come from zero, and the "something" it comes from is definitely more of the data you train on than the training method you choose.
Re: Publishers want billions, not millions, from AI
#5Presumably OpenAI and others (for the most part) are checking the licenses for content they use in training? Let’s say for example that they trained on the content of the entire archive of the New York Times… isn’t it safe to say they’d have purchased a license for that content from NYT? Wouldn’t “commercial use” cover this? Seems like the publishers just want a piece of the money because they want a piece of the mon…
Re: Publishers want billions, not millions, from AI
#6I am out of sympathy for big publishers. They have been greedy and predatory for decades, at the expense of their own content.
Re: Publishers want billions, not millions, from AI
#7Now that one company closed it off, is trying to regulate things, and is profiting from it: Let the lawyers win with fees.
Re: Publishers want billions, not millions, from AI
#8Presumably OpenAI and others (for the most part) are checking the licenses for content they use in training? Let’s say for example that they trained on the content of the entire archive of the New York Times… isn’t it safe to say they’d have purchased a license for that content from NYT? Wouldn’t “commercial use” cover this? Seems like the publishers just want a piece of the money because they want a piece of the mon…
The problem is that companies see that they undersold their content relative to the value it’s providing LLMs, and now want to renegotiate past and future deals with AI companies. Exactly as you said: they just want a bigger slice of pie.
Re: Publishers want billions, not millions, from AI
#9Re: Publishers want billions, not millions, from AI
#10LLMs are a genuinely exciting technology, but I am worried that a part of what they enable will turn out to be a sort of "content laundering." If you train your AI on a bunch of someone elses work, you can make it produce something very simiilar. You can extract whole aspects of an artist or writer's style and replicate their ideas much more easily now, and because it is generated inside this black box you have a sort of deniability when it comes to accusations of copying.