Live data from Hacker News

Are large language models a threat to digital public goods?

arxiv.org

21–30 of 163 posts

Re: Are large language models a threat to digital public goods?

#21

Earlier quoted context omitted.

And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?

I don't have good answers. I have some high-level intuitions. One of them is that creation costs of information are fixed, while its usefulness is unbounded, so it doesn't make sense to try and reward creators for each access/view/use, in perpetuity. Secondly, there's a lot of information laundering going on - any random book I read carries between a few to few hundred references to prior written work. What I pay for…

I think that is a somewhat narrow view. Maybe to make the contrast sharper: Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms?

The new situation isn't the same as search as that wasn't there to hide information sources or to immediately convert information into useful things (texts, guides, etc.).

Re: Are large language models a threat to digital public goods?

#22

Earlier quoted context omitted.

And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?

I don't have good answers. I have some high-level intuitions. One of them is that creation costs of information are fixed, while its usefulness is unbounded, so it doesn't make sense to try and reward creators for each access/view/use, in perpetuity. Secondly, there's a lot of information laundering going on - any random book I read carries between a few to few hundred references to prior written work. What I pay for…

There is a big potential role for open source or more specifically copyleft / free AI here that is released as a community project but can be monetized as well. The evidence from software is that there is lots of interest in contributing to such products.

Re: Are large language models a threat to digital public goods?

#23
post #18

Earlier quoted context omitted.

And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?

> Even if it isn't monetizeable IP, how to share the costs? The internet started thanks to ample government funding for research. So have many other technologies, including AI. I wonder if there's a way we could all somehow pool our resources and use that to pay for common goods that we all use. What would we call such a scheme?

We could definitely fund things. Or tax LLM firms in some broad way to redistribute to the respective societies for using their creations.

Re: Are large language models a threat to digital public goods?

#24

Earlier quoted context omitted.

I don't have good answers. I have some high-level intuitions. One of them is that creation costs of information are fixed, while its usefulness is unbounded, so it doesn't make sense to try and reward creators for each access/view/use, in perpetuity. Secondly, there's a lot of information laundering going on - any random book I read carries between a few to few hundred references to prior written work. What I pay for…

I think that is a somewhat narrow view. Maybe to make the contrast sharper: Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms? The new situation isn't the same as search as that wasn't there to hide information sources or to immediately convert information into useful things (texts, guides, etc.).

> Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms?

If this matters to you, then you shouldn't. But to flip this around: why should you care?

Unless you're doing some unique work targeting a global audience, the point when LLM gets trained on what you created is way outside space you'd normally care about. Trying to capture all the value your work generates does not lead to a good world.

Or maybe it's me who isn't profit-minded enough, but e.g. a lot of what I wrote on-line, including blog articles and commentary on Reddit and HN, has been used by search engines for free for a long time (over a decade, in some cases), and now is (most likely) part of the training corpora for LLMs. But I never believed, and still don't believe, that I'm entitled to some share of the gains LLMs (or search engines) make.

Re: Are large language models a threat to digital public goods?

#25

Earlier quoted context omitted.

Interesting, so we chuck in exabytes and more of data generated each day and then what?

The focus should eventually be on building AI such that they can be given access to data and then independently figure out new and creative ways to use it, rather than requiring humans to figure that out for them and then narrowly define: do this thing with the data I've provided. Given the scale of the data, they'll be better suited to that approach if it's breakthrough outcomes that we want.

Yes, maybe good to put the ideas around "dataism" (don't know a better word, sorry) really to the test.

Re: Are large language models a threat to digital public goods?

#26

Earlier quoted context omitted.

For reasons I can't articulate I see LLMs as a vehicle for removing the creators from their ideas. This is very different than search engines. If a search engine generates traffic for documented ideas it creates a community. An LLM based internet seems to remove the creator and shim itself in between for the sake of business.

It's tricky, because in many ways, this is achieving exactly what I, as user, want computers to do for me: give me information I requested, and only that. I very much do not care about who discovered/created/published it, except only if it helps me quickly ascertain the trustworthiness of said information. I do not want to be forced or prodded to establish relationships with creators or communities. I do not want the…

I don't know that I have a position yet. The problem I have is when a tech wizard says the tech is so complex nobody knows what is going on is never good.

>I do not want to be forced or prodded to establish relationships with creators or communities. I do not want their ads and upsells.

I feel the same way I lurked on HN for 7 years before signing up. I enjoy the lurking aspects of the internet.

I usually think the internet can survive anything.

This seems like the same disruption that Uber promoted to avoid regulation. Imagine all the great ideas we can generate and claim they are original to the unauditable AI thought process.

The last worst thing to happen to the internet IMHO was SEO. I sort think a schism may develop to the reality of the broadcast world and the narrowcast world.

I definitely think that valid uses exist.

Re: Are large language models a threat to digital public goods?

#27

This is a very intriguing study as it would appear the fact that LLMs need public data to train on, its incentive on the marketplace of ideas is to reduce the very thing that gives it its power. It contains the seeds of its own destruction, destroying the web and open data ethos, and yet another data point pointing at another AI winter.

For reasons I can't articulate I see LLMs as a vehicle for removing the creators from their ideas. This is very different than search engines. If a search engine generates traffic for documented ideas it creates a community. An LLM based internet seems to remove the creator and shim itself in between for the sake of business.

I've been wondering whether an LLM could list its sources, if the training data included source data for each document (perhaps in the form "The source of the following text is XYZ:")

Re: Are large language models a threat to digital public goods?

#29

They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…

I'd expect user engagement to be going down in the same trend that has been occurring over the past five years as Stackoverflow is less and less useful.

Re: Are large language models a threat to digital public goods?

#30

Earlier quoted context omitted.

I think that is a somewhat narrow view. Maybe to make the contrast sharper: Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms? The new situation isn't the same as search as that wasn't there to hide information sources or to immediately convert information into useful things (texts, guides, etc.).

> Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms? If this matters to you, then you shouldn't. But to flip this around: why should you care? Unless you're doing some unique work targeting a global audience, the point when LLM gets trained on what you created is way outside space you'd normally care about. Trying to capture all the value your work generates…

This isn't so much about compensation, but why should I help enrich a large, even more direct rent seeker?

Valuable information in a way is becoming more valuable for the LLM provider, so I would expect a drop in high value information in the public domain.

Post reply on HN