Live data from Hacker News

Are large language models a threat to digital public goods?

arxiv.org

141–150 of 163 posts

Re: Are large language models a threat to digital public goods?

#141

Earlier quoted context omitted.

I think that is somewhat off topic. I don't see why rent seeking via IP laws should be bad while doing it via provision of AI wouldn't be.

LLMs aren't just a mere database containing indexed copies of other peoples' IP. AI companies are charging you for access to a sophisticated automated reasoning system, that necessarily had to memorize half of the Internet in the process of becoming capable of (some approximation of) reasoning. (BTW. that you can even make a system this way is a huge breakthrough that's not being talked about enough.) But even if the…

I think we are talking past each other. Let me try to narrow down where I think we disagree.

1) LLM providers harvest a common to create their product (don't think we disagree here much).

2) What happens next is where we diverge, I suspect: I think they will use their products to extract rents from that common while you think they will provide a fairly priced service.

Ultimately time will tell how the business model shakes out. Both could even be happening in sequence.

Re: Are large language models a threat to digital public goods?

#142

Earlier quoted context omitted.

> And this is why humanity is going down the tubes... On the contrary - this is exactly how and why humanity built a technological civilization in the first place. Note that I didn't say > because you want something, you derive value from what you want, and yet you do not care about giving something back to who makes it. Yes, because it would be backward and limiting to do that. Note: I never said I don't want to giv…

What does your example have to do with this situation? The people in the bread supply chain get paid, the author of content we're discussing will never get paid by anyone, will never even get a bit if personal satisfaction from their analytics knowing last month x thousand people read that page and it hopefully helped them. It's completely zero reward, even worse it's completely zero feedback of any kind! This really…

I'm really struck how this is the first time I have seen a disincentive to freely share information on the internet.

I can't tell if I have aged out of some ideal or is it that the individual creative efforts are being homogenized into pseudo answers for someone to sell.

Re: Are large language models a threat to digital public goods?

#143

Earlier quoted context omitted.

For years people have been making travel blogs based on where they've visited and the practical information they've discovered, like experiences of visiting attractions or good places to stay in cities or how they got from one place to another. They monetised with ads and affiliate links so they could travel more based on that income. In LLM land, they get no monetisation any more because nobody visits their sites, i…

> They monetised with ads and affiliate links I consider this to be a problem on its own, but it's not relevant here because: > In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found. That can't possibly be true, because if it were, there wouldn't be any travel blogs anymore today. All that travel spam has been made redundant approxim…

> As for "A LOT of the useful information" - nope, can't think of a single case where ad/affiliate-supported site was a good information source, vs. just displacing a better free source.

https://stingynomads.com/annapurna-circuit-cost-planning/

Solid introduction and primer for the Annapurna Circuit, full of useful information from people who did it which has been kept up to date.

https://stingynomads.com/who-are-we/

> Today stingynomads.com is our full-time business and main source of income.

Now please tell me in what way is this not an example of an ad/affiliate supported site that provides a lot of useful information and what non ad/affiliate based resource has it displaced that was better? Cause I'm doubting someone would write up a better guide than that, publish it and not monetise it.

Re: Are large language models a threat to digital public goods?

#144

Earlier quoted context omitted.

That is true I see no reason obvious reason why the LL companies take pride in not being able to document ideation process. I have no justification but I feel it is deceitful not technical reasoning.

The issue here is that memorization of any distinguishable part of IP is an incidental aspect - those models aren't memorizing stuff, they're learning it. We don't expect people to keep track of the source of every single piece of information they encounter. It would arguably make learning impossible - as much for humans as for LLMs. As an intuition pump, when I write "2+2 = " and you mentally complete it with "4", s…

What is the hard technical barrier that makes the tracking of attribution for input sequences for LLM training impossible? I don't see any.

Re: Are large language models a threat to digital public goods?

#145

They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…

What happens after LLMs kill off SO and then seek more updated training data? It seems like these “deaths” are either temporary or something else will pop-up that continuously improves and trains an LLMs with proprietary data. There’s def a shift in the social contract of the internet. We’re shifting from “publish a few things you know a lot bit in exchange to read stuff other smart people have shared” to “solve a pr…

With programming atleast there can a validation step at LLM based system can deploy. Check if the suggested changes work before spitting out the answer.

Re: Are large language models a threat to digital public goods?

#146

Earlier quoted context omitted.

LLMs aren't just a mere database containing indexed copies of other peoples' IP. AI companies are charging you for access to a sophisticated automated reasoning system, that necessarily had to memorize half of the Internet in the process of becoming capable of (some approximation of) reasoning. (BTW. that you can even make a system this way is a huge breakthrough that's not being talked about enough.) But even if the…

I think we are talking past each other. Let me try to narrow down where I think we disagree. 1) LLM providers harvest a common to create their product (don't think we disagree here much). 2) What happens next is where we diverge, I suspect: I think they will use their products to extract rents from that common while you think they will provide a fairly priced service. Ultimately time will tell how the business model…

> 2) What happens next is where we diverge, I suspect: I think they will use their products to extract rents from that common while you think they will provide a fairly priced service.

Phrased like this, I can't really disagree with you. I don't expect a business to play fair in general, when it has a profitable option to do otherwise.

I guess my objection is more that right now, I don't see LLMs creating any kind of disincentive to publish quality content. In my eyes, LLMs are not a substitute for quality content in the first place - I see them more like using quality content to create a tool that competes with ad-hoc and shitty content.

That's not to say LLMs won't be able to eventually provide high-quality information on their own - but at that point, we'll have more important problems to deal with, such as chunk of humanity being rendered obsolete.

Re: Are large language models a threat to digital public goods?

#147

Earlier quoted context omitted.

I think that is somewhat off topic. I don't see why rent seeking via IP laws should be bad while doing it via provision of AI wouldn't be.

LLMs aren't just a mere database containing indexed copies of other peoples' IP. AI companies are charging you for access to a sophisticated automated reasoning system, that necessarily had to memorize half of the Internet in the process of becoming capable of (some approximation of) reasoning. (BTW. that you can even make a system this way is a huge breakthrough that's not being talked about enough.) But even if the…

LLMs are not reasoning systems. That's one of the major problems with them.

Re: Are large language models a threat to digital public goods?

#148
post #54

They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. "Second, we investigate whether ChatGPT is simply displ…

> They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. The moderation team and community in general on Stack…

I've arrived from google on "closed" questions I needed answers for, XY questions where I had question X but not secret question Y, people trying to XY answer extremely simple and straightforward X-and-only-X questions, questions unanswered for years with several upvotes and me-toos.

But shallow questions where the official docs are too raw or are missing a few specifics? Stack Overflow was good for that before chatgpt.

Re: Are large language models a threat to digital public goods?

#149

Earlier quoted context omitted.

The issue here is that memorization of any distinguishable part of IP is an incidental aspect - those models aren't memorizing stuff, they're learning it. We don't expect people to keep track of the source of every single piece of information they encounter. It would arguably make learning impossible - as much for humans as for LLMs. As an intuition pump, when I write "2+2 = " and you mentally complete it with "4", s…

What is the hard technical barrier that makes the tracking of attribution for input sequences for LLM training impossible? I don't see any.

When you make an omelette, what is the technical barrier making it practically impossible to tell which egg contributed how much to any given part of the meal?

It's roughly the same thing.

Re: Are large language models a threat to digital public goods?

#150

Earlier quoted context omitted.

Why not? I could see two ways it could work. First, it seems possible that if sources were in the training data like I described, then understanding of sources could be an emergent capability, just because the LLM reads "the source of the following is X." Second, maybe a trainer LLM could be tasked with reading the trainee's answers and any sources it provides, and judging whether the source is correct. But I'm no ex…

Well, you can train the LLM to "provide source", but LLMs are prone to hallucination. You run into the same problem with a trainer model; the trainer also has no way to confirm where the model actually got the answer from. One thing that may work is a fundamental architectural shift where the LLM looks up all its info as it needs it, and then you can just list the sources it actually used. Microsoft tried that with B…

Sure but I don't think it's necessary to confirm where the AI got something. There might be lots of sources. I often tell someone some fact I know, mention a source if I remember it, and possibly google a link. An AI could do much the same. In training, the trainer could just check whether the claimed sources actually say something similar to what the trainee claimed it said.

In operation, the "trainer" could do the same thing in the background. And then of course, human users could also check up on the sources if they need to be sure of catching hallucinations.

Post reply on HN