Are large language models a threat to digital public goods?
71–80 of 163 posts
Re: Are large language models a threat to digital public goods?
#72They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…
It seems like these “deaths” are either temporary or something else will pop-up that continuously improves and trains an LLMs with proprietary data.
There’s def a shift in the social contract of the internet. We’re shifting from “publish a few things you know a lot bit in exchange to read stuff other smart people have shared” to “solve a problem with an LLM and train it in a manner that can be helpful to others”
What’s unclear is how the middle of that will be monetized. Search engines figured out how to do it for the “publish and share” paradigm and made a boatload of money off of it after paying for the massive infrastructure required to index everything. How will LLMs do it without killing the data that trains it?
Re: Are large language models a threat to digital public goods?
#73Re: Are large language models a threat to digital public goods?
#74I won't upload any of my writing to the web anymore until this is all sorted out. You can make fun of me all you like, but it's taken me decades to get good at this, and I'll be damned if some soft-skinned SV kid with a MacBook uses my work to power his mill.
Re: Are large language models a threat to digital public goods?
#75Re: Are large language models a threat to digital public goods?
#76I see this as a rough parallel to "is the printing press a threat to illuminated manuscripts". Maybe it is under a very narrow view, but overall, improved propagation of ideas has led to improved dissemination of ideas, and it will this time too. People who's narrow world has been disrupted will perform all sorts of mental gymnastics to tell us how we're going to be worse off for it, but we won't. Ironically, interne…
Re: Are large language models a threat to digital public goods?
#77I see this as a rough parallel to "is the printing press a threat to illuminated manuscripts". Maybe it is under a very narrow view, but overall, improved propagation of ideas has led to improved dissemination of ideas, and it will this time too. People who's narrow world has been disrupted will perform all sorts of mental gymnastics to tell us how we're going to be worse off for it, but we won't. Ironically, interne…
Even then you'd want attribution to survive.
Eliminating or changing the incentives people have for doing things is the absolutely surest way to change behavior.
Postulating that all published work can be appropriated at will by certain private for-profit entities is the death knell of the knowledge economy as we know it.
Re: Are large language models a threat to digital public goods?
#78Earlier quoted context omitted.
The problem with this sort of comparison is that up until now, new ideas came from humans and technology advances merely helped to spread ideas, or greased the wheels. Now you can generate new ideas. Often without any skill of your own. This is bad at first glance because look what happened when smart phones became popular: no one remembers phone numbers of their family, people can’t remember directions when driving.…
> when smart phones became popular: no one remembers phone numbers of their family, people can’t remember directions when driving. This is a concern, though I'd argue unrelated to the concern about people no longer contributing publicly online, and one that was already present with SO. I've seen SO posts specifically reference or criticize the "copy paste" crowd that is just taking the answer and putting it into thei…
Re: Are large language models a threat to digital public goods?
#79Earlier quoted context omitted.
> Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms? If this matters to you, then you shouldn't. But to flip this around: why should you care? Unless you're doing some unique work targeting a global audience, the point when LLM gets trained on what you created is way outside space you'd normally care about. Trying to capture all the value your work generates…
A lot of people justifiably care because making that information is their livelihood. Your entitlement to the labor of others is gross.
Re: Are large language models a threat to digital public goods?
#80They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…
What happens after LLMs kill off SO and then seek more updated training data? It seems like these “deaths” are either temporary or something else will pop-up that continuously improves and trains an LLMs with proprietary data. There’s def a shift in the social contract of the internet. We’re shifting from “publish a few things you know a lot bit in exchange to read stuff other smart people have shared” to “solve a pr…