Live data from Hacker News

Are large language models a threat to digital public goods?

arxiv.org

81–90 of 163 posts

Re: Are large language models a threat to digital public goods?

#81
post #51

I won't upload any of my writing to the web anymore until this is all sorted out. You can make fun of me all you like, but it's taken me decades to get good at this, and I'll be damned if some soft-skinned SV kid with a MacBook uses my work to power his mill.

I get the feeling.

When I felt like writing again after years of closing my previous blog, it though about it for a while before committing to https://bitecode.dev.

But eventually, I realized that

- I also write for myself, not just for others.

- People like reading things without having to prompt for it. So they will read the blog because it's nice and topics come to them even if they don't think that they need to know.

- ChatGPT doesn't have opinion. It tries very hard to be balanced. Your blog will have an opinion.

- There is more to the experience you provide than just knowledge. You can add tools, exercises, videos, etc. Which GPT cannot, for now, replicate.

- GPT can replicate style, but will not by default. People will come for your style as well. And pics. And design. And jokes.

- People value the interaction they feel when they content seems like there is a person behind it. They get attached. They develop sympathy.

- A blog puts things in context. "If you want to know that, you probably need to know that". It also gives information about what happen right now.

- Humans curate. In a world where creating crappy content is very cheap, a good filter has tremendous value.

So yes, you will be scanned, and replicated. Doesn't mean you don't have value in writing what you do.

And making something of value is nice.

This may change in 10 years or so. Maybe the LLM will be able to do all that. But depriving yourself of a rewarding experience right now for the fear of what might happen is not worth it.

Re: Are large language models a threat to digital public goods?

#82

They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…

What happens after LLMs kill off SO and then seek more updated training data? It seems like these “deaths” are either temporary or something else will pop-up that continuously improves and trains an LLMs with proprietary data. There’s def a shift in the social contract of the internet. We’re shifting from “publish a few things you know a lot bit in exchange to read stuff other smart people have shared” to “solve a pr…

> “publish a few things you know a lot bit in exchange to read stuff other smart people have shared

I feel like this hasn't been the majority of internet users experience for ~20 years

Re: Are large language models a threat to digital public goods?

#83

They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…

To expect change in vote counts and/or ratios, you have to assume that users have sort of "vote budget" that they are going to spend one way or another. If voting is mostly independent variable and mostly depends on particular user and particular post then you should not expect voting patterns/counts to change meaningfully.

> Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be.

Given the later's facility with the language, I would expect it to be a better search engine. I would expect that ChatGPT is replacing Google as sort of tokenizer. I.e. search pipeline changes from "form natural question -> input to google -> go to SO -> if first few links do not yield answer post new question" to "form natural question -> input to chatgpt -> extract keyword tokens -> input to google -> go to SO".

There is an important bit in the article:

> > Using data on programming language popularity on GitHub, we find that the most widely used languages tend to have larger relative declines in posting activity.

Here we can form a hypothesis that reduction in post frequency comes from entry level posts with posters not knowing what to search for. Under this hypothesis ChatGPT has strongest effect on users using SO as knowledge base rather than Q&A forum. This user type distinction would affect posting frequency much more heavily than voting patterns.

Re: Are large language models a threat to digital public goods?

#84

Earlier quoted context omitted.

I have no doubt the last thing the LLM firms want is to attribute their sources. I always see the claim, heck we don't know where the ideas come from that is impossible.

Yes it’s just an engineering problem, there is no a priori reason not to do it.

That's a very big "just" given that there is no good way to train for this.

Re: Are large language models a threat to digital public goods?

#85

They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…

I'd expect user engagement to be going down in the same trend that has been occurring over the past five years as Stackoverflow is less and less useful.

As someone who is getting their start with coding, where/what forums exist that have a high quality/helpful community? My biggest struggle has been with relatively simple questions – with a broad stroked theme/issue to 'em for the most part. AKA having a mentor or just a group of coders who are willing to help out if you are willing to be an active member (but I'm relatively useless, aka active maybe in an off-topic lounge part of it). I'd appreciate it if ya got the sauce, a PM if you'd like to keep it on the DL perhaps? Thanks!

Re: Are large language models a threat to digital public goods?

#86

Earlier quoted context omitted.

What happens after LLMs kill off SO and then seek more updated training data? It seems like these “deaths” are either temporary or something else will pop-up that continuously improves and trains an LLMs with proprietary data. There’s def a shift in the social contract of the internet. We’re shifting from “publish a few things you know a lot bit in exchange to read stuff other smart people have shared” to “solve a pr…

Would be cool to have projects write documentation in a form that's annotated for LLMs, ie. premarked topic breaks.

> have projects write documentation ...

I kind of expect those projects would want ChatGPT to write the documentation. ;)

Re: Are large language models a threat to digital public goods?

#87

I see this as a rough parallel to "is the printing press a threat to illuminated manuscripts". Maybe it is under a very narrow view, but overall, improved propagation of ideas has led to improved dissemination of ideas, and it will this time too. People who's narrow world has been disrupted will perform all sorts of mental gymnastics to tell us how we're going to be worse off for it, but we won't. Ironically, interne…

> improved propagation of ideas has led to improved dissemination of ideas, and it will this time too

I sincerely hope you are shitposting. So called Web 2.0, the platformized web has placed huge incentives to stifle propagation and dissemination of ideas. However, these incentives are somewhat distributed under human review, therefore there can be bubbles with opposing ideas. LLMs centralize censorship, by design.

Currently you can find "BrawndoLounge" and "ToiletWater" subreddits that heavily censor opposing ideas. Even if one is full of hate speech or whatever, there still are opposing ideas there. If LLM guardians decide that one of these sources is "toxic", in an LLM-led web that viewpoint simply vanishes.

Re: Are large language models a threat to digital public goods?

#88

They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…

I'd expect user engagement to be going down in the same trend that has been occurring over the past five years as Stackoverflow is less and less useful.

They used a difference-in-differences[1] to control for such longitudinal/temporal effects. Since ChatGPT isn't readily available in China or Russia, but StackOverflow is, they can compare how SO changed pre and post ChatGPT in countries where it is widely available and compare that to the pre and post in countries where it isn't available (basically as the control).

[1] https://en.wikipedia.org/wiki/Difference_in_differences

Re: Are large language models a threat to digital public goods?

#89
post #60

Earlier quoted context omitted.

> Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms? If this matters to you, then you shouldn't. But to flip this around: why should you care? Unless you're doing some unique work targeting a global audience, the point when LLM gets trained on what you created is way outside space you'd normally care about. Trying to capture all the value your work generates…

A lot of people justifiably care because making that information is their livelihood. Your entitlement to the labor of others is gross.

Ideas are copied by reading or hearing them. You can't own your ideas now, unless by own you mean horde. The perpetual creators rights you want extended are already artificial and require a non-trivial amount of our GDP to enforce and they still stifle future creation in a lot of areas.

Most people are paid for doing things every day, they don't get to create one thing and never work again. Expanding creators compensation laws is regressive and only helps a few elites survive job uncertainly, not the bulk of the people. We're better off limiting this sort of thing specifically to help everyone advance, share the knowledge.

Re: Are large language models a threat to digital public goods?

#90

Earlier quoted context omitted.

> when smart phones became popular: no one remembers phone numbers of their family, people can’t remember directions when driving. This is a concern, though I'd argue unrelated to the concern about people no longer contributing publicly online, and one that was already present with SO. I've seen SO posts specifically reference or criticize the "copy paste" crowd that is just taking the answer and putting it into thei…

Why is that a concern? When’s the last time you needed to remember a phone number or directions for travel?

Because of the implied effect on things like attention span, memory, and cognition in general. It's the same reason that mastering fundamental mathematics is critical, even though we have calculators. Trends in the Flynn Effect are relevant here. [1] The Flynn Effect is the observation that "real" IQ values were increasing, dramatically, for decades. In recent decades (since births in ~1970), this effect has not only decreased but reversed in the West. And it cannot be explained by dysgenics alone, since it even shows up in same-family cohorts. That suggests environmental causes are likely playing some role.

Those coming to adulthood in the 90s were the first generation to really get to experience mass, endless, and cheap dopamine driven digital entertainment. Notably the [positive] Flynn effect is still in full swing in places like China and India, where digital entertainment market saturation has taken longer. So they'll effectively work as perfect experiments. If it reverses over the coming decades, as everybody over in these places now has their heads shoved in e.g. smart phones as much as anywhere else, it should be telling.

[1] - https://www.sciencealert.com/iq-scores-falling-in-worrying-r...

Post reply on HN