Live data from Hacker News

GenAI, the snake eating its own tail

ybrikman.com

21–30 of 138 posts

Re: GenAI, the snake eating its own tail

#21

The article feels very confused to me. Example 1 is bad, StackOverflow had clearly plateaued and was well into the downward freefall by the time ChatGPT was released. Example 2 is apparently "open source" but it's actually just Tailwind which unfortunately had a very susceptible business model. And I don't really think the framing here that it's eating its own tail makes sense. It's also confusing to me why they're t…

> They can try to solve that problem Well, they could always try actually paying content creators. Unlike - for instance - StackOverflow.

StackOverflow as built back in the days of Web 2.0 where the idea was that user generated content formed in the days of the (relatively) altruistic web.

There isn't any clean way to do "contributor gets paid" without adding in an entire mess of "ok, where is the money coming from? Paywalls? Advertising? Subscriptions?" and then also get into the mess of international money transfers (how do you pay someone in Iran from the US?)

And then add in the "ok, now the company is holding payment information of everyone(?) ..." and data breaches and account hacking is now so much more of an issue.

Once you add money to it, the financial inceptives and gamification collide to make it simply awful.

Re: GenAI, the snake eating its own tail

#22

“We can’t put the genie back in the bottle.” Actually we can. And we will.

How?

Only way I could see it is if there's enough pushback on them taking everyone's power and water (and computer parts) in a world where power and water are becoming increasingly unstable. But I feel like defeating Ai because there is not enough consistent water and power to give them means there is more pressing issues at hand...

Re: GenAI, the snake eating its own tail

#23

“We can’t put the genie back in the bottle.” Actually we can. And we will.

Agreed, it's funny how people have taken unrestrained use of AI as an axiom at this point. There very much is still time to significantly control it + regulate it. Is there enough appetite by those in power (across the political spectrum)? Right now I don't think so.

>There very much is still time to significantly control it + regulate it.

There's also huge financial momentum shoving AI through the world's throat. Even if AI was proven to be a failure today, it would still be pushed for many years because of the momentum.

I just don't see how that can be reversed.

Re: GenAI, the snake eating its own tail

#24

The article feels very confused to me. Example 1 is bad, StackOverflow had clearly plateaued and was well into the downward freefall by the time ChatGPT was released. Example 2 is apparently "open source" but it's actually just Tailwind which unfortunately had a very susceptible business model. And I don't really think the framing here that it's eating its own tail makes sense. It's also confusing to me why they're t…

The proposed solution is also pretty confused: > For each response, the GenAI tool lists the sources from which it extracted that content, perhaps formatted as a list of links back to the content creators, sorted by relevance, similar to a search engine This literally isn’t possible given the architecture of transformer models and there’s no indication it will ever be.

Could you ELI5 why this isn't possible? Google's search result AI summary shows the links for example.

Re: GenAI, the snake eating its own tail

#26
post #3

We’ve seen decades of growing wage gaps and erosion of labors strength. The current elites don’t really care to enrich the people. Why would they care to do anything about this problem? They likely don’t see it as a problem at all. If they did actually stumble on AGI (assuming it didn’t eat them too) it would be used by a select few to enslave or remove the rest of us.

Not sure why this is being downvoted. It's spot on. You see folks like Dario et al. raising the alarm bells about what they claim is coming... while working as hard as they can to bring that gloomy future to fruition. No one in power is going to help unless there's money in it.

According to Trump, "If it was up to Stephen [Miller], there would only be 100 million people in this country — and all of them would look like him."

Re: GenAI, the snake eating its own tail

#27
GenAI changes the dynamics of information systems so fundamentally that our entire notion of intellectual property is being upended.

Copyright was predicated on the notion that ideas and styles can not be protected, but that explicit expressive works can. For example, a recipe can't be protected, but the story you wrap around it that tells how your grandma used to make it would be.

LLMs are particularly challenging to wrangle with because they perform language alchemy. They can (and do) re-express the core ideas, styles, themes, etc. without violating copyright.

People deem this 'theft' and 'stealing' because they are trying to reconcile the myth of intellectual property with reality, and are also simultaneously sensing the economic ladder being pulled up by elites who are watching and gaming the geopolitical world disorder.

There will be a new system of value capture that content creators need to position for, which is to be seen as a more valuable source of high quality materials than an LLM, serving a specific market, and effectively acquiring attention to owned properties and products.

It will not be pay-per-crawl. Or pay-per-use. It will be an attention game, just like everything in the modern economy.

Attention is the only way you can monetize information.

Re: GenAI, the snake eating its own tail

#28
post #5
post #2

Pay per crawl of StackOverflow wouldn't encourage me to post more on StackOverflow. (Not that I was anyway.) Presumably you'd need to pay content creators, but that seems quite inefficient: 1. I pay OpenAI 2. OpenAI rev shares to StackOverflow 3. StackOverflow mostly keeps that money, but shares some with me for posting 4. I get some money back to help pay OpenAI? This is nonsense. And if the frontier labs are right…

>If the argument is sustainability of training, I'm skeptical we need these payment models. That seems to be the argument: LLM adoption leads to drop of organic training data, leading LLMs to eventually plateau, and we'll be left without the user-generated content we relied on for a while (like SO) and with subpar LLM. That's what I'm getting from the article anyway.

There are so many things wrong with the points this article repeats, but those are soundbites at this point so I'm not sure one can even argue against them anymore.

Still, for the one about organic data (or "pre-war steel") drying out, it's not a threat to model development at all. People repeating this point don't realize that we already have way more data than we need. We got to where we are by brute-forcing the problem - throwing more data at a simple training process. If new "pristine" data were to stop flowing now, we still a) have decent pre-trained base models, and a dataset that's more than sufficient to train more of them, and b) lots of low-hanging fruits to pick in training approaches, architectures and data curation, that will allow to get more performance out of same base data.

That, and the fact that synthetic data turned out to be quite effective after all, especially in the latter phases of training. No surprise there, for many classes of problems this is how we learn as well. Anyone who has experience studying math for maturity exam / university entry exams knows this: the best way to learn is to solve lots of variations of the same set of problems. These variations are all synthetic data, until recently generated by hand, but even their trivial nature doesn't make them less effective at teaching.

Re: GenAI, the snake eating its own tail

#30
post #5

Earlier quoted context omitted.

>If the argument is sustainability of training, I'm skeptical we need these payment models. That seems to be the argument: LLM adoption leads to drop of organic training data, leading LLMs to eventually plateau, and we'll be left without the user-generated content we relied on for a while (like SO) and with subpar LLM. That's what I'm getting from the article anyway.

The article gets the part about organic data dying off right. Look at Google SERP's for an example. Almost nobody clicks through to the source anymore, so ad revenue is drying up for them and people are publishing less or publishing in places that pay them directly and live behind a paywall like Medium. Which means Google has less data to work with. That said, what it misses is that the AI prompts themselves become a…

>AI prompts themselves become a giant source of data.

Good point, but can it match the old organic data? I'm skeptical. For one, the LLM environment lacks any truth or consensus mechanism that the old SO-like sites had. 100s of users might have discussed the same/similar technical problem with an LLM, but there's no way (afaik) for the AI to promote good content and demote bad ones, as it (AI) doesn't have the concept of correctness/truth. Also, the old sites were two-sided, with humans asking _and_ answering questions, while they are only on the asking side with AI.

Post reply on HN