Live data from Hacker News

Publishers want billions, not millions, from AI

semafor.com

61–70 of 78 posts

Re: Publishers want billions, not millions, from AI

#61
post #53
post #17

Big media companies go fuck yourself. That nightmare scenario, for Levin, would turn a Food & Wine review into a simple text recommendation of a bottle of Malbec. If there is just one wine review on the web which recommends Malbec, AI will not start recommending it. If there are many such reviews, then yes, AI will tell you "Many reviews recommend Malbec.". Just like a human can tell you about what they learned, our…

That is not correct. LLMs have no inventory of their input datasets. If you give it a billion reviews recommending Malbec, it does not know that many reviews recommend Malbec. It simply has a very good comprehension of what a Malbec review looks like. In order to know that many reviews recommend Malbec, it needs to be taught this fact specifically, or infer it from context.

Prompt:

Do more Wikipedia pages mention the word "blue" or do more Wikipedia pages mention the word "Accismus"?

ChatGPT:

To determine whether more Wikipedia pages mention the word "blue" or the word "Accismus," we can use a simple search on the Wikipedia website.

Searching for the term "blue" on Wikipedia yields a large number of results, as the color blue is a common topic across various fields. On the other hand, "Accismus" is a relatively obscure term, and it is less likely to be mentioned in a significant number of Wikipedia pages.

Therefore, it is highly probable that more Wikipedia pages mention the word "blue" compared to the word "Accismus."

Re: Publishers want billions, not millions, from AI

#62

I am reminded somewhat of the ride & delivery apps here. While they did use tech to enable some cool things like demand pricing and efficient route planning, a big part of their innovation really came from using that tech to shovel most of the risk and cost of providing the service onto independent contractors. LLMs are a genuinely exciting technology, but I am worried that a part of what they enable will turn out to…

The real "deniability" of copying will come when the NLP community gets off of its collective rear end and implements actual prompt engineering (i.e. any technique I mention in here: https://gist.github.com/Hellisotherpeople/45c619ee22aac6865c...). Using "prompt blending" (i.e. https://github.com/ljleb/prompt-fusion-extension) will give genuine "deniability" on the grounds that "A beautiful story by {stephen_king|plato|aristotle|virgina_wolf}" will be very unique and moreover such kinds of interpolation are for the most part how real human writers write.

Re: Publishers want billions, not millions, from AI

#63

Earlier quoted context omitted.

>they should be paying out if they are training on data owned by someone Virtually everything you "know" is because of current and past humans. Have you been paying everyone for all of that? If you want to be a great author and you read books by great writers.. then become a successful writer, are you paying all those authors for training you? I realize there is a difference with being able to absorb bulk data and al…

But you do pay for all of that in todays society. Teachers are paid, tuition to schools is paid, feedback from tutors is paid, textbooks are paid for, copies of books are paid for, movies and TV are paid for. The experience you learn on the job is paid for via yours or your coworker’s salary. I get the argument that it’s not copying it’s learning, but it’s also very different than a human learning by observation. And…

The training is not the issue. It’s the reproduction. I expect long term the models will still be trained on copyrighted material, but it will refuse to recite any of it.

Re: Publishers want billions, not millions, from AI

#64

I am reminded somewhat of the ride & delivery apps here. While they did use tech to enable some cool things like demand pricing and efficient route planning, a big part of their innovation really came from using that tech to shovel most of the risk and cost of providing the service onto independent contractors. LLMs are a genuinely exciting technology, but I am worried that a part of what they enable will turn out to…

I'm a mod/admin on a topic site, similar to HN in structure... one of the posts I removed a few weeks ago literally made me feel ill... it was a guide on setting up a website, then using an LLM to generate hundreds of "articles" for said site. I've started seeing content sites that are obviously generated content looking at it... mostly in terms of recipe content for lower carb, or sugar free... some list ingredients…

My Twitter verse is filled with people making websites with thousands of pageviews from ai generated content. Google obviously is turning a blind eye to this possibly because they are planning on something a lot bigger than chatgpt with ai. And when it happens the internet will be littered with ai generated content that it is impossible for anyone to accuse them and them alone for stealing content.

Re: Publishers want billions, not millions, from AI

#65

I am reminded somewhat of the ride & delivery apps here. While they did use tech to enable some cool things like demand pricing and efficient route planning, a big part of their innovation really came from using that tech to shovel most of the risk and cost of providing the service onto independent contractors. LLMs are a genuinely exciting technology, but I am worried that a part of what they enable will turn out to…

I'm a mod/admin on a topic site, similar to HN in structure... one of the posts I removed a few weeks ago literally made me feel ill... it was a guide on setting up a website, then using an LLM to generate hundreds of "articles" for said site. I've started seeing content sites that are obviously generated content looking at it... mostly in terms of recipe content for lower carb, or sugar free... some list ingredients…

That's fascinating. I'm also seeing the sort of effects that this recession has been having on people especially online, where a lot of the guides you're talking about are being pushed more and more as a way to "side hustle". Having a side hustle is pushed more and more these days because a "main" hustle doesn't exist anymore due to wage stagnation.

Part of the popularity of these sort of things stems entirely from how much money is reportedly up for grabs. If you tell people they can do this and bring in 2k a month on the side, they're all in. Likewise, if we had enough funds to not have to worry, these sort of things would rarely be used as there's no real joy in the pursuit of it. My theory on why AI is accelerating garbage content much faster than any other potential benefits.

Even the highlight of recipes is notable, recipe blogs are great earners due to the stability of ads and output of content. This becomes a target for AI (when I say AI here I mean people consuming AI API) due to the high ad value target. That way even if they do eventually drop, get found out, reported to google, all of that, they would have still made enough to do it again.

Re: Publishers want billions, not millions, from AI

#66
post #2

Presumably OpenAI and others (for the most part) are checking the licenses for content they use in training? Let’s say for example that they trained on the content of the entire archive of the New York Times… isn’t it safe to say they’d have purchased a license for that content from NYT? Wouldn’t “commercial use” cover this? Seems like the publishers just want a piece of the money because they want a piece of the mon…

>Seems like the publishers just want a piece of the money because they want a piece of the money.

Sure, but of course, we all know the actual value came from the writers.

Re: Publishers want billions, not millions, from AI

#67
post #17

Big media companies go fuck yourself. That nightmare scenario, for Levin, would turn a Food & Wine review into a simple text recommendation of a bottle of Malbec. If there is just one wine review on the web which recommends Malbec, AI will not start recommending it. If there are many such reviews, then yes, AI will tell you "Many reviews recommend Malbec.". Just like a human can tell you about what they learned, our…

>Nobody should be allowed to ... extort money from society for writing wine reviews.

So how do you propose to pay for someone to write wine reviews?

Re: Publishers want billions, not millions, from AI

#68

Back when it was OpenAI, I easily could side with the open source software. Now that one company closed it off, is trying to regulate things, and is profiting from it: Let the lawyers win with fees.

Well, "open"ai opened pandora's box. Let the games begin.

Re: Publishers want billions, not millions, from AI

#69

Earlier quoted context omitted.

Taxi drivers already were independent contractors.

And taxi drivers themselves caused Uber to be so popular. The amount of times I've gotten into a cab in London before Uber and the cabbie was like: "Cash only". What's that machine for then? "Credit card machine down". Yea right... I absolutely hate what Uber did with 'contractors' but I cannot deny the fact that I like being able to travel for work without having to worry about whether the cab has a credit card mach…

> And taxi drivers themselves caused Uber to be so popular.

That entirely depends on where you are. Where I am, ride-share services were never better than taxis in any way except for price.

Re: Publishers want billions, not millions, from AI

#70
post #61
post #53

Earlier quoted context omitted.

That is not correct. LLMs have no inventory of their input datasets. If you give it a billion reviews recommending Malbec, it does not know that many reviews recommend Malbec. It simply has a very good comprehension of what a Malbec review looks like. In order to know that many reviews recommend Malbec, it needs to be taught this fact specifically, or infer it from context.

Prompt: Do more Wikipedia pages mention the word "blue" or do more Wikipedia pages mention the word "Accismus"? ChatGPT: To determine whether more Wikipedia pages mention the word "blue" or the word "Accismus," we can use a simple search on the Wikipedia website. Searching for the term "blue" on Wikipedia yields a large number of results, as the color blue is a common topic across various fields. On the other hand, "…

That is not relevant to what I said. The model can have knowledge about its corpus or tooling to ask questions about it.

But the training process is not a library it can draw upon or answer questions about intrinsically.

Post reply on HN