Live data from Hacker News

Are large language models a threat to digital public goods?

arxiv.org

121–130 of 163 posts

Re: Are large language models a threat to digital public goods?

#121
post #51

I won't upload any of my writing to the web anymore until this is all sorted out. You can make fun of me all you like, but it's taken me decades to get good at this, and I'll be damned if some soft-skinned SV kid with a MacBook uses my work to power his mill.

> You can make fun of me all you like

Those that do make fun of you are victim blaming. A lot of these folks stating that if you put out in the public then it's not yours anymore sound like criminals to be fair.

Re: Are large language models a threat to digital public goods?

#122
post #51

I won't upload any of my writing to the web anymore until this is all sorted out. You can make fun of me all you like, but it's taken me decades to get good at this, and I'll be damned if some soft-skinned SV kid with a MacBook uses my work to power his mill.

> You can make fun of me all you like Those that do make fun of you are victim blaming. A lot of these folks stating that if you put out in the public then it's not yours anymore sound like criminals to be fair.

Or people who don't make anything. It's very easy to be generous with other people's work.

Re: Are large language models a threat to digital public goods?

#123
post #51

I won't upload any of my writing to the web anymore until this is all sorted out. You can make fun of me all you like, but it's taken me decades to get good at this, and I'll be damned if some soft-skinned SV kid with a MacBook uses my work to power his mill.

Sometimes I'm thinking about publishing AI generated content that is clearly labeled as such. Just to make the anyone who scrapes it to train their model a little bit worse.

Re: Are large language models a threat to digital public goods?

#124

Earlier quoted context omitted.

> You can make fun of me all you like Those that do make fun of you are victim blaming. A lot of these folks stating that if you put out in the public then it's not yours anymore sound like criminals to be fair.

Or people who don't make anything. It's very easy to be generous with other people's work.

Techno communism essentially. "From each according to his ability, to each according to his needs". And just like that type of economic setup a handful of bros will reap the benefits - in this case sam altman's clique and others like them.

Re: Are large language models a threat to digital public goods?

#125

Earlier quoted context omitted.

> And this is why humanity is going down the tubes... On the contrary - this is exactly how and why humanity built a technological civilization in the first place. Note that I didn't say > because you want something, you derive value from what you want, and yet you do not care about giving something back to who makes it. Yes, because it would be backward and limiting to do that. Note: I never said I don't want to giv…

Regarding the baker example, some form of compensation is eventually directed to the employees, the farmers etc. even though you could say there are many layers of indirection. In case of LLMs, no compensation is directed to the person authoring the information. While it may not be a problem for the consumer of the information, it removes any incentives for the people authoring the information to continue doing so, w…

Depends what the incentives are. Some folks like the community- that doesn't change. A website like honestwargamer or goonhammer has a clan. Mostly because they are insightful and friendly. But if an AI scraped their content, it would have to build a better goonhammer? Maybe. But I bet it would have half a dozen nerds contributing and tinkering at the backend, painting the new minis that came out, talking about tournament results etc.. That is current, constant fresh relevant content, based on IRL activity. Very hard to replicate. For honestwargamer... Build a better twitch stream... that seems even less likely. they run a round and stream games, review results, collate stats. And then engage with the audience in twitch.

So when you say "contribute to the internet", this is what I consume and.. I'm sure there are similar examples in every niche- fishing, golf, coding, AI art creation ...

No I don't see this as gloomy scenario, and content creators- the goonhammers and honestwargamers, creatives, are still going to get paid (a bit, they were never rich), maybe in new ways.

Re: Are large language models a threat to digital public goods?

#126
post #82

Earlier quoted context omitted.

What happens after LLMs kill off SO and then seek more updated training data? It seems like these “deaths” are either temporary or something else will pop-up that continuously improves and trains an LLMs with proprietary data. There’s def a shift in the social contract of the internet. We’re shifting from “publish a few things you know a lot bit in exchange to read stuff other smart people have shared” to “solve a pr…

> “publish a few things you know a lot bit in exchange to read stuff other smart people have shared I feel like this hasn't been the majority of internet users experience for ~20 years

It was never the experience of "majority". Majority doesn't create.

It was, nevertheless, internet.

Re: Are large language models a threat to digital public goods?

#127
post #123
post #51

I won't upload any of my writing to the web anymore until this is all sorted out. You can make fun of me all you like, but it's taken me decades to get good at this, and I'll be damned if some soft-skinned SV kid with a MacBook uses my work to power his mill.

Sometimes I'm thinking about publishing AI generated content that is clearly labeled as such. Just to make the anyone who scrapes it to train their model a little bit worse.

And Abimelech fought against the city all that day; and he took the city, and slew the people that were therein: and he beat down the city, and sowed it with salt.

Re: Are large language models a threat to digital public goods?

#128

Earlier quoted context omitted.

Yes it’s just an engineering problem, there is no a priori reason not to do it.

That's a very big "just" given that there is no good way to train for this.

Why not? I could see two ways it could work.

First, it seems possible that if sources were in the training data like I described, then understanding of sources could be an emergent capability, just because the LLM reads "the source of the following is X."

Second, maybe a trainer LLM could be tasked with reading the trainee's answers and any sources it provides, and judging whether the source is correct.

But I'm no expert, hence my question.

Re: Are large language models a threat to digital public goods?

#129

Earlier quoted context omitted.

> And this is why humanity is going down the tubes... On the contrary - this is exactly how and why humanity built a technological civilization in the first place. Note that I didn't say > because you want something, you derive value from what you want, and yet you do not care about giving something back to who makes it. Yes, because it would be backward and limiting to do that. Note: I never said I don't want to giv…

What does your example have to do with this situation? The people in the bread supply chain get paid, the author of content we're discussing will never get paid by anyone, will never even get a bit if personal satisfaction from their analytics knowing last month x thousand people read that page and it hopefully helped them. It's completely zero reward, even worse it's completely zero feedback of any kind! This really…

> What does your example have to do with this situation?

It's addressing GP's complaint about me pointing out the indirect and transactional nature of the interaction between information producer and consumer.

> the author of content we're discussing will never get paid by anyone

That's... not my problem? Bear with me here.

> will never even get a bit if personal satisfaction from their analytics knowing last month x thousand people read that page and it hopefully helped them.

Aha!

So we're talking specifically about content creators that publish for free in hopes of maximizing a number on their analytics? That's healthy neither for them nor the society at large. Or you mean people publishing content for free to make money off ads? Yeah, I don't mind that content to disappear entirely.

Note that outside of web publishing, it was never the expectation of an author to have any idea how many people read their work, much less get paid for every single "read event". They only got a lump sum or a fraction of first sale of a printed work - but had no insight or control over further circulation of parts of entirety of their works. Being able to resell your books or magazines, or give or lend it to friends, or borrow some from a library, are all good things.

All that was true before AI, and people found reasons to write new books, or to publish quality content on-line, for free and without advertising or telemetry. LLMs don't change that. If anything, they may reduce readership, not publication.

> I think much information will retreat to places like Discord or locked down login only versions of sites like Stackoverflow.

This has already been happening for the past couple years; LLMs, again, don't change anything here.

Re: Are large language models a threat to digital public goods?

#130

Earlier quoted context omitted.

That's a very big "just" given that there is no good way to train for this.

Why not? I could see two ways it could work. First, it seems possible that if sources were in the training data like I described, then understanding of sources could be an emergent capability, just because the LLM reads "the source of the following is X." Second, maybe a trainer LLM could be tasked with reading the trainee's answers and any sources it provides, and judging whether the source is correct. But I'm no ex…

Well, you can train the LLM to "provide source", but LLMs are prone to hallucination. You run into the same problem with a trainer model; the trainer also has no way to confirm where the model actually got the answer from. One thing that may work is a fundamental architectural shift where the LLM looks up all its info as it needs it, and then you can just list the sources it actually used. Microsoft tried that with Bing, but it turns out the model will search for a website, read it, and then ignore what the website says and claim "according to this website, ". So it's definitely not easy at any rate.
Post reply on HN