Live data from Hacker News

Are large language models a threat to digital public goods?

arxiv.org

41–50 of 163 posts

Re: Are large language models a threat to digital public goods?

#41

Earlier quoted context omitted.

It's tricky, because in many ways, this is achieving exactly what I, as user, want computers to do for me: give me information I requested, and only that. I very much do not care about who discovered/created/published it, except only if it helps me quickly ascertain the trustworthiness of said information. I do not want to be forced or prodded to establish relationships with creators or communities. I do not want the…

And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?

I'm organizing the publishing of my thoughts of technology development and design. This has been something I have been mulling for decades. Completely uncompensated. I'm not doing any of this for reasons I can understand it is just what I think about and do.

Originally I saw a website as a way to hang out a shingle. Until recently I was thinking I could just publish away and maybe someone would hire me based off the website.

Currently I don't feel the same way regarding publishing on the web. I will be me more guarded in what I share.

Re: Are large language models a threat to digital public goods?

#42

Earlier quoted context omitted.

It's tricky, because in many ways, this is achieving exactly what I, as user, want computers to do for me: give me information I requested, and only that. I very much do not care about who discovered/created/published it, except only if it helps me quickly ascertain the trustworthiness of said information. I do not want to be forced or prodded to establish relationships with creators or communities. I do not want the…

And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?

I want attribution if I inspire a thought in AI.

I'm surprised the nerve the issue has struck in me.

Re: Are large language models a threat to digital public goods?

#43

I see this as a rough parallel to "is the printing press a threat to illuminated manuscripts". Maybe it is under a very narrow view, but overall, improved propagation of ideas has led to improved dissemination of ideas, and it will this time too. People who's narrow world has been disrupted will perform all sorts of mental gymnastics to tell us how we're going to be worse off for it, but we won't. Ironically, interne…

The problem with this sort of comparison is that up until now, new ideas came from humans and technology advances merely helped to spread ideas, or greased the wheels. Now you can generate new ideas. Often without any skill of your own. This is bad at first glance because look what happened when smart phones became popular: no one remembers phone numbers of their family, people can’t remember directions when driving.…

> when smart phones became popular: no one remembers phone numbers of their family, people can’t remember directions when driving.

This is a concern, though I'd argue unrelated to the concern about people no longer contributing publicly online, and one that was already present with SO. I've seen SO posts specifically reference or criticize the "copy paste" crowd that is just taking the answer and putting it into their code without thought. Some people will use any technology blindly.

Re: Are large language models a threat to digital public goods?

#46

Earlier quoted context omitted.

And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?

I want attribution if I inspire a thought in AI. I'm surprised the nerve the issue has struck in me.

It isn’t clear to me (other than it’s an open engineering problem) why LLs couldn’t also include attribution as part of training. Also tracking attribution could lead to some insights on how its internal representations in vector space are created.

Re: Are large language models a threat to digital public goods?

#47
post #27

Earlier quoted context omitted.

I've been wondering whether an LLM could list its sources, if the training data included source data for each document (perhaps in the form "The source of the following text is XYZ:")

I have no doubt the last thing the LLM firms want is to attribute their sources. I always see the claim, heck we don't know where the ideas come from that is impossible.

Yes it’s just an engineering problem, there is no a priori reason not to do it.

Re: Are large language models a threat to digital public goods?

#48

Earlier quoted context omitted.

I want attribution if I inspire a thought in AI. I'm surprised the nerve the issue has struck in me.

It isn’t clear to me (other than it’s an open engineering problem) why LLs couldn’t also include attribution as part of training. Also tracking attribution could lead to some insights on how its internal representations in vector space are created.

That is true I see no reason obvious reason why the LL companies take pride in not being able to document ideation process. I have no justification but I feel it is deceitful not technical reasoning.

Re: Are large language models a threat to digital public goods?

#50

I see this as a rough parallel to "is the printing press a threat to illuminated manuscripts". Maybe it is under a very narrow view, but overall, improved propagation of ideas has led to improved dissemination of ideas, and it will this time too. People who's narrow world has been disrupted will perform all sorts of mental gymnastics to tell us how we're going to be worse off for it, but we won't. Ironically, interne…

The problem with this sort of comparison is that up until now, new ideas came from humans and technology advances merely helped to spread ideas, or greased the wheels. Now you can generate new ideas. Often without any skill of your own. This is bad at first glance because look what happened when smart phones became popular: no one remembers phone numbers of their family, people can’t remember directions when driving.…

Skill generating ideas is over rated.

Implementation is when skill matters.

I can imagine post scarcity Star Trek like ideas but not implement them. My skills at CE and SWE allow me to implement reasonable solutions there.

Chasing ideas for the sake of chasing ideas is a form of bike shedding and premature optimization.

Post reply on HN