Live data from Hacker News

Are large language models a threat to digital public goods?

arxiv.org

51–60 of 163 posts

Re: Are large language models a threat to digital public goods?

#52

I see this as a rough parallel to "is the printing press a threat to illuminated manuscripts". Maybe it is under a very narrow view, but overall, improved propagation of ideas has led to improved dissemination of ideas, and it will this time too. People who's narrow world has been disrupted will perform all sorts of mental gymnastics to tell us how we're going to be worse off for it, but we won't. Ironically, interne…

> improved propagation of ideas has led to improved dissemination of ideas

Prove this.

Re: Are large language models a threat to digital public goods?

#54

They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…

> They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this.

The moderation team and community in general on Stack Overflow is so toxic I'm not even sure you could control for that effect well enough to arrive at this conclusion. I would argue people are leaving because it's easier to ask ChatGPT your question than be flamed and banned for asking how to do something. A half right answer from ChatGPT is better than getting marked duplicate and closed because the moron moderation team can't detect nuance.

Re: Are large language models a threat to digital public goods?

#55

I see this as a rough parallel to "is the printing press a threat to illuminated manuscripts". Maybe it is under a very narrow view, but overall, improved propagation of ideas has led to improved dissemination of ideas, and it will this time too. People who's narrow world has been disrupted will perform all sorts of mental gymnastics to tell us how we're going to be worse off for it, but we won't. Ironically, interne…

There is a weird sort of analogy where universities spent the last 20 years putting their classes online, or associate professors and grad students uploaded YouTube deep-dive videos, etc. Point is, enrollment at Universities are way down now and you could probably draw some kind of link between the two. However, there are also links between jobs that no longer care, and GenZ is a smaller cohort, and also maybe more o…

There's also this "enrolment cliff" the universities have been worried about for a long time. Apparently the pool of students is itself shrinking.

https://web.archive.org/web/20230315152647/https://www.forbe...

Re: Are large language models a threat to digital public goods?

#56

Earlier quoted context omitted.

I have no doubt the last thing the LLM firms want is to attribute their sources. I always see the claim, heck we don't know where the ideas come from that is impossible.

Yes it’s just an engineering problem, there is no a priori reason not to do it.

I'm not sure it's all that easy though. We don't entirely know how LLMs do some of the things they do, and we can't interpret what's in them. They don't internally look up particular sources, it's just a big mess of connection weights.

Maybe my simple scheme would be all it needs. Or maybe it needs some new breakthrough and right now nobody knows how to do it. I was hoping some resident expert could let me know.

Re: Are large language models a threat to digital public goods?

#57
post #51

I won't upload any of my writing to the web anymore until this is all sorted out. You can make fun of me all you like, but it's taken me decades to get good at this, and I'll be damned if some soft-skinned SV kid with a MacBook uses my work to power his mill.

Only correct take imo.

Re: Are large language models a threat to digital public goods?

#58

This is a very intriguing study as it would appear the fact that LLMs need public data to train on, its incentive on the marketplace of ideas is to reduce the very thing that gives it its power. It contains the seeds of its own destruction, destroying the web and open data ethos, and yet another data point pointing at another AI winter.

Literal tragedy of the commons. Overexploiting this resource will inevitably diminish it. Lack of human data to sample from will make models worse.

Re: Are large language models a threat to digital public goods?

#59

This is a very intriguing study as it would appear the fact that LLMs need public data to train on, its incentive on the marketplace of ideas is to reduce the very thing that gives it its power. It contains the seeds of its own destruction, destroying the web and open data ethos, and yet another data point pointing at another AI winter.

> its incentive on the marketplace of ideas is to reduce the very thing that gives it its power

It's not clear what the effect will be of LLMs consuming more and more of their own output and/or waste products.

One possibility might be a feedback system that leads to superintelligence which eventually becomes incomprehensible to humans.

Another might be a feedback system that leads to increased garbage output that eventually devolves into incomprehensible noise and nonsense.

The two extremes might not be easily distinguishable.

Re: Are large language models a threat to digital public goods?

#60

Earlier quoted context omitted.

I think that is a somewhat narrow view. Maybe to make the contrast sharper: Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms? The new situation isn't the same as search as that wasn't there to hide information sources or to immediately convert information into useful things (texts, guides, etc.).

> Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms? If this matters to you, then you shouldn't. But to flip this around: why should you care? Unless you're doing some unique work targeting a global audience, the point when LLM gets trained on what you created is way outside space you'd normally care about. Trying to capture all the value your work generates…

A lot of people justifiably care because making that information is their livelihood. Your entitlement to the labor of others is gross.
Post reply on HN