Live data from Hacker News

Are large language models a threat to digital public goods?

arxiv.org

61–70 of 163 posts

Re: Are large language models a threat to digital public goods?

#61

Earlier quoted context omitted.

Interesting, so we chuck in exabytes and more of data generated each day and then what?

The focus should eventually be on building AI such that they can be given access to data and then independently figure out new and creative ways to use it, rather than requiring humans to figure that out for them and then narrowly define: do this thing with the data I've provided. Given the scale of the data, they'll be better suited to that approach if it's breakthrough outcomes that we want.

No AI we currently have can do this.

Re: Are large language models a threat to digital public goods?

#62

Earlier quoted context omitted.

And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?

I don't have good answers. I have some high-level intuitions. One of them is that creation costs of information are fixed, while its usefulness is unbounded, so it doesn't make sense to try and reward creators for each access/view/use, in perpetuity. Secondly, there's a lot of information laundering going on - any random book I read carries between a few to few hundred references to prior written work. What I pay for…

> One of them is that creation costs of information are fixed, while its usefulness is unbounded, so it doesn't make sense to try and reward creators for each access/view/use, in perpetuity.

The word "creation" is loaded. No one "creates" content. They discover it hidden in some idea-space... occasionally even two people might discover the same thing. The same melody, the same verse of a poem, the same fragment of art.

The idea that one should be rewarded, but the other is slandered the infringer is amusingly dumb.

> but an LLM generating me a recipe based on associations created from being trained on millions of recipes, this feels like it should be in the clear, at least from user's POV

But how will we entrench the rent-seekers?

Re: Are large language models a threat to digital public goods?

#63
>> models like ChatGPT efficiently provide users with information about various topics

>> But since users interact privately with the model, these models may drastically reduce the amount of publicly available human-generated data and knowledge resources.

The premise seems a stretch when the models are simply providing a better way to find the information. It is essentially reducing the number of searches or low quality questions floating around.

For creator, it is just an aid to improve the creation. The models help them find duplicates better or simulate variations faster.

This paper appears to construct a case of walled garden by mis-representing user questions as the actual content/creation.

User search questions always remained private with the site owners (Google search, Bing etc.,)

Re: Are large language models a threat to digital public goods?

#64
post #51

I won't upload any of my writing to the web anymore until this is all sorted out. You can make fun of me all you like, but it's taken me decades to get good at this, and I'll be damned if some soft-skinned SV kid with a MacBook uses my work to power his mill.

I agree with you. Programmers at these AI companies basically create a wealth concentration mechanism that diverts the money resulting from the value of our work into their pockets.

Re: Are large language models a threat to digital public goods?

#65

Earlier quoted context omitted.

For reasons I can't articulate I see LLMs as a vehicle for removing the creators from their ideas. This is very different than search engines. If a search engine generates traffic for documented ideas it creates a community. An LLM based internet seems to remove the creator and shim itself in between for the sake of business.

It's tricky, because in many ways, this is achieving exactly what I, as user, want computers to do for me: give me information I requested, and only that. I very much do not care about who discovered/created/published it, except only if it helps me quickly ascertain the trustworthiness of said information. I do not want to be forced or prodded to establish relationships with creators or communities. I do not want the…

> I very much do not care about who discovered/created/published it, except only if it helps me quickly ascertain the trustworthiness of said information. I do not want to be forced or prodded to establish relationships with creators or communities. I do not want their ads and upsells.

And this is why humanity is going down the tubes...because you want something, you derive value from what you want, and yet you do not care about giving something back to who makes it.

Re: Are large language models a threat to digital public goods?

#66

Earlier quoted context omitted.

The focus should eventually be on building AI such that they can be given access to data and then independently figure out new and creative ways to use it, rather than requiring humans to figure that out for them and then narrowly define: do this thing with the data I've provided. Given the scale of the data, they'll be better suited to that approach if it's breakthrough outcomes that we want.

Yes, maybe good to put the ideas around "dataism" (don't know a better word, sorry) really to the test.

People have been testing them for ages. There's something there, but not nearly as much as the hype implies. At the same time, the hype seems to only increase.

Anyway, I really like that name. Non-dataist AIs are clearly the best kind.

Re: Are large language models a threat to digital public goods?

#67

>> models like ChatGPT efficiently provide users with information about various topics >> But since users interact privately with the model, these models may drastically reduce the amount of publicly available human-generated data and knowledge resources. The premise seems a stretch when the models are simply providing a better way to find the information. It is essentially reducing the number of searches or low qual…

I don't know about that. Firstly, do we know that the site can financially withstand a gigantic drop in traffic gracefully enough to avoid becoming the next Quora? Advertising doesn't pay out more for developers with novel technical questions. Secondly, will a failed ChatGPT inquiry still lead to an SO post so new good quality questions keep rolling in? What if they start to lose prominence in search rankings because they wouldn't be the go-to anymore? I could see people using reddit or the like to serve the same purpose in a much friendlier format.

Re: Are large language models a threat to digital public goods?

#68
post #51

I won't upload any of my writing to the web anymore until this is all sorted out. You can make fun of me all you like, but it's taken me decades to get good at this, and I'll be damned if some soft-skinned SV kid with a MacBook uses my work to power his mill.

He could've read your writing and done it anyway.

Re: Are large language models a threat to digital public goods?

#69
post #3

Does this mean we can’t adopt new languages since ChatGPT is frozen to 2022 coding knowledge?

I assume at some point they'll start a regular update cycle, or even continuous training. It's not going to be frozen in 2022 forever.

Re: Are large language models a threat to digital public goods?

#70

This is a very intriguing study as it would appear the fact that LLMs need public data to train on, its incentive on the marketplace of ideas is to reduce the very thing that gives it its power. It contains the seeds of its own destruction, destroying the web and open data ethos, and yet another data point pointing at another AI winter.

For reasons I can't articulate I see LLMs as a vehicle for removing the creators from their ideas. This is very different than search engines. If a search engine generates traffic for documented ideas it creates a community. An LLM based internet seems to remove the creator and shim itself in between for the sake of business.

This is the only valid take against LLM's I've seen to date.

At the same time, could it not be construed as humanity leveling up? We've abstracted away that, now low, level of thought. Much like we don't have people pressing the button in the elevator. Most developers and writers are still better than an LLM, but its good enough to replace their input on simple tasks.

Writing is being commoditized much like other industries. The moment you release a new vacuum cleaner there's ten others that do the same thing and nobody is fussed over who invented it. People still know who to go for a premium vacuum though.

Post reply on HN