Live data from Hacker News

Are large language models a threat to digital public goods?

arxiv.org

31–40 of 163 posts

Re: Are large language models a threat to digital public goods?

#31

They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…

the hypothesis is wrong to begin with

they should ask people that ask questions, not people passively looking for answers

the experience asking a question on stackexchange sites is horrible! each tag is its own community with edicts and customs you have no idea about which derail your path to actually getting an answer, the auto-moderation system is completely broken because it thinks your question should be replaced with one from 2013, or the human moderators unilaterally fix your question which then causes the auto-moderator to think its similar to the one from 2013 and unceremoniously replaces it

With chatgpt you dont need to prove anything about what you tried, you dont need a reputation score to do anything, you dont need to be a steward of every upvote for all eternity, you dont need to engage in meta discussion to get your post out of deletion, and it answers faster. You don't need to be told to ask a separate question if you have any followups, with all of the same gamble of problems, while chatgpt already assumes what you will have a followup question about and just tells you all the gotchas. Because its read all the forum posts and has seen what the recurring issues are.

its better, faster, and cheaper time wise

its not a direct comparison, chatgpt is more like a pair programmer. If stackoverflow had an adhoc pair programming live session you could hop into it would be a better comparison.

Re: Are large language models a threat to digital public goods?

#33
post #18

Earlier quoted context omitted.

And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?

> Even if it isn't monetizeable IP, how to share the costs? The internet started thanks to ample government funding for research. So have many other technologies, including AI. I wonder if there's a way we could all somehow pool our resources and use that to pay for common goods that we all use. What would we call such a scheme?

> I wonder if there's a way we could all somehow pool our resources and use that to pay for common goods that we all use. What would we call such a scheme?

Is this tongue in cheek? I think it's called the government and taxes! :)

Re: Are large language models a threat to digital public goods?

#34

Earlier quoted context omitted.

And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?

I don't have good answers. I have some high-level intuitions. One of them is that creation costs of information are fixed, while its usefulness is unbounded, so it doesn't make sense to try and reward creators for each access/view/use, in perpetuity. Secondly, there's a lot of information laundering going on - any random book I read carries between a few to few hundred references to prior written work. What I pay for…

The biggest reason I care/have fear about information I will share freely getting absorbed into AI is...

I definitely believe that the I could get sued for my own ideas. The chance of lawyer claiming AI say it is their idea sounds horrible.

Re: Are large language models a threat to digital public goods?

#35
post #27

Earlier quoted context omitted.

For reasons I can't articulate I see LLMs as a vehicle for removing the creators from their ideas. This is very different than search engines. If a search engine generates traffic for documented ideas it creates a community. An LLM based internet seems to remove the creator and shim itself in between for the sake of business.

I've been wondering whether an LLM could list its sources, if the training data included source data for each document (perhaps in the form "The source of the following text is XYZ:")

I have no doubt the last thing the LLM firms want is to attribute their sources. I always see the claim, heck we don't know where the ideas come from that is impossible.

Re: Are large language models a threat to digital public goods?

#36

I see this as a rough parallel to "is the printing press a threat to illuminated manuscripts". Maybe it is under a very narrow view, but overall, improved propagation of ideas has led to improved dissemination of ideas, and it will this time too. People who's narrow world has been disrupted will perform all sorts of mental gymnastics to tell us how we're going to be worse off for it, but we won't. Ironically, interne…

The problem with this sort of comparison is that up until now, new ideas came from humans and technology advances merely helped to spread ideas, or greased the wheels. Now you can generate new ideas. Often without any skill of your own.

This is bad at first glance because look what happened when smart phones became popular: no one remembers phone numbers of their family, people can’t remember directions when driving. Now you don’t even need the ability to generate ideas.

And it’s not clear to me that LLMs democratises information. LLM have guard rails, are potentially trained on copyrighted data. Need to be run by large companies…

Re: Are large language models a threat to digital public goods?

#37

I see this as a rough parallel to "is the printing press a threat to illuminated manuscripts". Maybe it is under a very narrow view, but overall, improved propagation of ideas has led to improved dissemination of ideas, and it will this time too. People who's narrow world has been disrupted will perform all sorts of mental gymnastics to tell us how we're going to be worse off for it, but we won't. Ironically, interne…

There is a weird sort of analogy where universities spent the last 20 years putting their classes online, or associate professors and grad students uploaded YouTube deep-dive videos, etc. Point is, enrollment at Universities are way down now and you could probably draw some kind of link between the two.

However, there are also links between jobs that no longer care, and GenZ is a smaller cohort, and also maybe more of an interest in trades as a path forward.

I'm not suggesting the paper linked here is right or wrong, but we should keep an open mind about all the other potential reasons why things suddenly shift. One random reason could be that during this overly oppressive year of layoffs engineers have either stopped needing to search or haven't had time to work on a project that needed it.

Also there's smarter IDEs which are likely the #1 reason I've stopped searching the internet for one-off small things like "what's the API to deep copy an array in Rust." Documented autocomplete has gotten really good and things like copilot are just icing on the cake.

Re: Are large language models a threat to digital public goods?

#38

They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…

the hypothesis is wrong to begin with they should ask people that ask questions, not people passively looking for answers the experience asking a question on stackexchange sites is horrible! each tag is its own community with edicts and customs you have no idea about which derail your path to actually getting an answer, the auto-moderation system is completely broken because it thinks your question should be replaced…

This is a great point. I think the research is "biased" to begin with, in that it feels like you could easily p-hack your way into a "chat gpt has degraded X" kind of paper if that's what you want to show. But with SO in particular, you're right that there are actually clear improvements that it makes as a competitor.

I can't think of a way, buy it would be interesting to see if there is a similar effect on Wikipedia contributions.

Re: Are large language models a threat to digital public goods?

#39

Earlier quoted context omitted.

> Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms? If this matters to you, then you shouldn't. But to flip this around: why should you care? Unless you're doing some unique work targeting a global audience, the point when LLM gets trained on what you created is way outside space you'd normally care about. Trying to capture all the value your work generates…

This isn't so much about compensation, but why should I help enrich a large, even more direct rent seeker? Valuable information in a way is becoming more valuable for the LLM provider, so I would expect a drop in high value information in the public domain.

Perhaps there will be a drop in high value information in the public domain, but right now, I can't exactly see LLMs impacting the incentives for creation and sharing of that information. I don't see how LLMs would make someone go "oh well, AI is here, I might as well stop providing people with no-strings-attached high quality information", if the existence of search engines didn't make them stop already.

Re: Are large language models a threat to digital public goods?

#40
post #27

Earlier quoted context omitted.

I've been wondering whether an LLM could list its sources, if the training data included source data for each document (perhaps in the form "The source of the following text is XYZ:")

I have no doubt the last thing the LLM firms want is to attribute their sources. I always see the claim, heck we don't know where the ideas come from that is impossible.

Perhaps, but EleutherAI for example is a nonprofit that has trained several open source LLMs.
Post reply on HN