So this was the outcome I guessed would need to happen, but I didn't think publishers would be the ones to push for it, but I guess they are. The thing is if this upsets you on some gut level, this is just a sign for how valuable information to train on is. May be billions is too much, but hundreds of millions? If your immediate thought is "this must be stopped, this isn't fair" it is a signal of how reliant AI at al…
Publishers want billions, not millions, from AI
11–20 of 78 posts
Re: Publishers want billions, not millions, from AI
#12They are in no position to make demands. The current data that has been scraped for AI models is probably good enough to be able to generate synthetic data.
Even if it's not, what are publishers going to do about people scraping their content? Absolutely nothing.
Trying to force this issue feels like a good way to be disintermediated quicker.
Re: Publishers want billions, not millions, from AI
#13Presumably OpenAI and others (for the most part) are checking the licenses for content they use in training? Let’s say for example that they trained on the content of the entire archive of the New York Times… isn’t it safe to say they’d have purchased a license for that content from NYT? Wouldn’t “commercial use” cover this? Seems like the publishers just want a piece of the money because they want a piece of the mon…
Ha.
There is a difficult to describe culture in the ML community, from the top researchers to casual tinkerers... Everything is bleeding edge and moving at blinding speed. Some devs don't even stop to think about things like licensing or easy reproducibility/packaging, because they are too busy moving onto the next thing.
This thread on how a popular repo was unlicensed and violating other licenses for months is a good example: https://github.com/AUTOMATIC1111/stable-diffusion-webui/issu...
OpenAI comes from this culture, even if they are a more commercial company now.
Re: Publishers want billions, not millions, from AI
#14Hilarious and desperate. They are in no position to make demands. The current data that has been scraped for AI models is probably good enough to be able to generate synthetic data. Even if it's not, what are publishers going to do about people scraping their content? Absolutely nothing. Trying to force this issue feels like a good way to be disintermediated quicker .
Re: Publishers want billions, not millions, from AI
#15I am reminded somewhat of the ride & delivery apps here. While they did use tech to enable some cool things like demand pricing and efficient route planning, a big part of their innovation really came from using that tech to shovel most of the risk and cost of providing the service onto independent contractors. LLMs are a genuinely exciting technology, but I am worried that a part of what they enable will turn out to…
As a blogger of 10+ years this has been very interesting to realize. I can use ChatGPT to write really good first drafts because “in the style of Swizec Teller” works as a prompt. The results are way better than the usual generic corporate drone style of output.
BUT! And this is an important but. The insights it produces remain at that generic corporate drivel level. Even with a full list of bullet points of cool insights, an LLM takes those and averages them out into a big pile of meh. Every sharp insight gets dulled by “AI protections”, every hot take is two-sided to death, etc.
Despite superficially following my style, the AI seems unable to catch your attention. Everything it writes is a little boring, a little too long, a dash tiring to read. The spark is missing.
Also my style has evolved since 2021 and that’s missing.
Re: Publishers want billions, not millions, from AI
#16I am out of sympathy for big publishers. They have been greedy and predatory for decades, at the expense of their own content.
Yes, I agree, they could have invested in the various AI companies, instead the kept bleeding their content creators of all they have. Well the boat sailed, see you later big publishers.
Not really what I meant.
> instead the kept bleeding their content creators of all they have
But yes, that. And their customers too.
Re: Publishers want billions, not millions, from AI
#17 That nightmare scenario, for Levin,
would turn a Food & Wine review into
a simple text recommendation of a
bottle of Malbec.
If there is just one wine review on the web which recommends Malbec, AI will not start recommending it. If there are many such reviews, then yes, AI will tell you "Many reviews recommend Malbec.". Just like a human can tell you about what they learned, our artificial brothers have the same right to learn.Nobody should be allowed to bend the law and extort money from society for writing wine reviews.
Re: Publishers want billions, not millions, from AI
#18Big media companies go fuck yourself. That nightmare scenario, for Levin, would turn a Food & Wine review into a simple text recommendation of a bottle of Malbec. If there is just one wine review on the web which recommends Malbec, AI will not start recommending it. If there are many such reviews, then yes, AI will tell you "Many reviews recommend Malbec.". Just like a human can tell you about what they learned, our…
Re: Publishers want billions, not millions, from AI
#19I am reminded somewhat of the ride & delivery apps here. While they did use tech to enable some cool things like demand pricing and efficient route planning, a big part of their innovation really came from using that tech to shovel most of the risk and cost of providing the service onto independent contractors. LLMs are a genuinely exciting technology, but I am worried that a part of what they enable will turn out to…
Re: Publishers want billions, not millions, from AI
#20Presumably OpenAI and others (for the most part) are checking the licenses for content they use in training? Let’s say for example that they trained on the content of the entire archive of the New York Times… isn’t it safe to say they’d have purchased a license for that content from NYT? Wouldn’t “commercial use” cover this? Seems like the publishers just want a piece of the money because they want a piece of the mon…
OpenAI absolutely does not have licenses for 99%+ of the content used to build their models. They're following the standard tech company model of "negotiate forgiveness rather than ask for permission".