Live data from Hacker News

Are large language models a threat to digital public goods?

arxiv.org

131–140 of 163 posts

Re: Are large language models a threat to digital public goods?

#131
post #60

Earlier quoted context omitted.

A lot of people justifiably care because making that information is their livelihood. Your entitlement to the labor of others is gross.

Ideas are copied by reading or hearing them. You can't own your ideas now, unless by own you mean horde. The perpetual creators rights you want extended are already artificial and require a non-trivial amount of our GDP to enforce and they still stifle future creation in a lot of areas. Most people are paid for doing things every day, they don't get to create one thing and never work again. Expanding creators compens…

I think that is somewhat off topic. I don't see why rent seeking via IP laws should be bad while doing it via provision of AI wouldn't be.

Re: Are large language models a threat to digital public goods?

#132

Earlier quoted context omitted.

Or people who don't make anything. It's very easy to be generous with other people's work.

Techno communism essentially. "From each according to his ability, to each according to his needs". And just like that type of economic setup a handful of bros will reap the benefits - in this case sam altman's clique and others like them.

Private companies abusing works shared in good faith for free in order to profit is like communism? I'm sorry, that's just absurd "anything I don't like is communism"-level thinking. These are organizations operating under the incentives of capitalism to achieve the goal of capitalism (and lots of tech enthusiasts trying to come up with post-hoc justifications for the shiny new toy).

Re: Are large language models a threat to digital public goods?

#133

Earlier quoted context omitted.

Perhaps there will be a drop in high value information in the public domain, but right now, I can't exactly see LLMs impacting the incentives for creation and sharing of that information. I don't see how LLMs would make someone go "oh well, AI is here, I might as well stop providing people with no-strings-attached high quality information", if the existence of search engines didn't make them stop already.

For years people have been making travel blogs based on where they've visited and the practical information they've discovered, like experiences of visiting attractions or good places to stay in cities or how they got from one place to another. They monetised with ads and affiliate links so they could travel more based on that income. In LLM land, they get no monetisation any more because nobody visits their sites, i…

> They monetised with ads and affiliate links

I consider this to be a problem on its own, but it's not relevant here because:

> In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found.

That can't possibly be true, because if it were, there wouldn't be any travel blogs anymore today. All that travel spam has been made redundant approximately around the time Flickr was created, and every interesting location ever has been photographed from every interesting angle in a way neither me, nor you, nor your favorite travel blogger could ever hope to match. All the information they post has also been posted many times over by travel bloggers that came before.

The point being: travel information and photography is worthless commodity these days. Travel bloggers (or Instagrammers, or whatever) are not in the business of selling information. They're selling dreams and personal experiences. The photos and information are necessary as delivery vector ("social object"), but by themselves are worthless and not the point. The point is entertainment, and ideally getting you trapped in a parasocial relationship with the travel blogger/grammer, which gives them a recurring revenue stream.

> In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found.

It's the same model as with most other ad-monetized social media publishing. People will keep visiting them for the same reason they visit them now, and for the same reason they have their favorite youtubers and tiktokers. LLMs and other generative models don't change anything here, at least not short-to-mid-term, because they can't convincingly replicate human connection and keep it up for long.

(Also, I personally don't buy that travel instagrammers can actually sustain their travel lifestyle through ads and affiliate marketing and sponsorship deals. I suspect most are funded in some way, whether by family wealth or by services performed while traveling around.)

> The search engines actively supported these authors, by sending them people who needed the answers they had. (...) A LOT of the useful information on the web was built on similar feedback loops and they go away in LLM land.

Hard disagree. The only feedback loop this created in practice is the one that displaces quality information from the Internet - the combination of SEO and ad-based monetization means the most scummy players are the ones with most money to stalk every conceivable search query. The results speak for themselves: making a Google query for pretty much any topic of interest to general population will give you only content marketing sites - results that carry negative knowledge, as in if you waste your time reading them, you'll come more misinformed about the topic than you were before. If LLMs make all that go away, I'm 100% for it.

As for "A LOT of the useful information" - nope, can't think of a single case where ad/affiliate-supported site was a good information source, vs. just displacing a better free source.

Re: Are large language models a threat to digital public goods?

#134
post #50

Earlier quoted context omitted.

Skill generating ideas is over rated. Implementation is when skill matters. I can imagine post scarcity Star Trek like ideas but not implement them. My skills at CE and SWE allow me to implement reasonable solutions there. Chasing ideas for the sake of chasing ideas is a form of bike shedding and premature optimization.

You can only say this right now where there are an abundance of ideas and less skill to implement them. But even implementation requires creative thinking and the ability to generate ideas. Now we want to take away the need for the first part. This is NOT like going from handwriting to printing or radio to TV. Something new is going on. I use ChatGPT for work in a limited way. If you haven’t tried it, I suggest that…

We could say this for a long time now; all the historical squabbles over religious story is evidence useless ideation is an innate human feature that wastes time/resources, and the advancement of skill at implementation is what’s actually improved our day to day.

Re: Are large language models a threat to digital public goods?

#135

Earlier quoted context omitted.

Then perhaps they should find other means of livelihood, instead of preventing the rest of the world from making full use of the information and technology available to it.

They will and less information will be put into places that are freely accessible. If it's put anywhere at all it'll be put behind login only/paywalled/unscrapable places that LLM's can't access.

Why would they suddenly paywall information if they weren't already?

The way I see it, there are roughly three groups of information providers:

1. Those who do it pro bono - because they feel like its a worthwhile thing to do, or because they believe in by "pay it forward", or otherwise because they haven't even thought that what they share is worth trying to extract rent from.

2. Those who do it "for free", as a way to lure people to where they can expose them to ads, affiliate marketing, upsells, or other such schemes - making money by being predators using information as bait.

3. Those who just put up a paywall, being up front that they're selling information, not giving it away.

(There's also a weird "in-between" group of publishers that are almost like 1., except they're being funded out of marketing budgets of companies that figure providing quality information is good advertising.)

LLMs don't change anything for group #1. They may compete with group #3, but that's business as usual, not anything transformative. The group that's directly affected is #2, which also happens to be the group that produces all the garbage on the Internet, so I'm actually very happy to see them forced to find a more useful way of making money. Since group #2 produces "information" that's arguably negative knowledge on the net, it's likely to improve the amount of quality information you'll be able to find on-line.

Re: Are large language models a threat to digital public goods?

#136
post #54

Earlier quoted context omitted.

> They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. The moderation team and community in general on Stack…

I have anecdotal stories about aa friend corroborating this. He had been rebuked and experienced an unwelcoming response on Stack Overflow in response to his questions while trying to learn web development, and found ChatGPT a great teacher in comparison.

I'm a very experienced developer. Sometimes I have questions that are easier to ask a person than go dig through miles of documentation. I could easily frame a question correctly to get a response but even I feel extremely unwelcome there. I couldn't imagine being junior.

Re: Are large language models a threat to digital public goods?

#137

Earlier quoted context omitted.

Ideas are copied by reading or hearing them. You can't own your ideas now, unless by own you mean horde. The perpetual creators rights you want extended are already artificial and require a non-trivial amount of our GDP to enforce and they still stifle future creation in a lot of areas. Most people are paid for doing things every day, they don't get to create one thing and never work again. Expanding creators compens…

I think that is somewhat off topic. I don't see why rent seeking via IP laws should be bad while doing it via provision of AI wouldn't be.

LLMs aren't just a mere database containing indexed copies of other peoples' IP. AI companies are charging you for access to a sophisticated automated reasoning system, that necessarily had to memorize half of the Internet in the process of becoming capable of (some approximation of) reasoning.

(BTW. that you can even make a system this way is a huge breakthrough that's not being talked about enough.)

But even if they were a mere database indexing copies of other peoples' IP, then - copyright issues notwithstanding - the de-bullshittifying of information retrieval process alone would be service worth paying a lot of money for.

Re: Are large language models a threat to digital public goods?

#138

Earlier quoted context omitted.

It isn’t clear to me (other than it’s an open engineering problem) why LLs couldn’t also include attribution as part of training. Also tracking attribution could lead to some insights on how its internal representations in vector space are created.

That is true I see no reason obvious reason why the LL companies take pride in not being able to document ideation process. I have no justification but I feel it is deceitful not technical reasoning.

The issue here is that memorization of any distinguishable part of IP is an incidental aspect - those models aren't memorizing stuff, they're learning it. We don't expect people to keep track of the source of every single piece of information they encounter. It would arguably make learning impossible - as much for humans as for LLMs.

As an intuition pump, when I write "2+2 = " and you mentally complete it with "4", should I chastise you for not completing it with "4, as per ${your elementary class math textbook} and ${that other book you read as a kid}, corroborated by ${your first math teacher} and ${your parent} quoting ${some other work}"?

Re: Are large language models a threat to digital public goods?

#139

Earlier quoted context omitted.

I have really enjoyed using https://phind.com , which includes attribution in its responses.

Phind’s base model, which is GPT3.5/4 doesn’t itself do attribution, it’s made to do that with prompt engineering which provides the most relevant materials on the web based on a word embedding vector search, and then asks it to reference each source in the answer.

I mean, this is more-less what a student does when writing a paper, when they're forced to cite their sources. They first come up with an idea based on their own understanding/recollection, then they try to figure out where did they first took that idea from. If they remember a specific source, they'll cite that; if they don't (because there may not be one specific source they learned from), they'll search for some existing work that expresses the idea in question, and cite that.

I.e. in case of both the student and an LLM, correct citation doesn't actually mean the idea originates from the cited work - only that the work contains this idea.

Re: Are large language models a threat to digital public goods?

#140

Earlier quoted context omitted.

What does your example have to do with this situation? The people in the bread supply chain get paid, the author of content we're discussing will never get paid by anyone, will never even get a bit if personal satisfaction from their analytics knowing last month x thousand people read that page and it hopefully helped them. It's completely zero reward, even worse it's completely zero feedback of any kind! This really…

> What does your example have to do with this situation? It's addressing GP's complaint about me pointing out the indirect and transactional nature of the interaction between information producer and consumer. > the author of content we're discussing will never get paid by anyone That's... not my problem? Bear with me here. > will never even get a bit if personal satisfaction from their analytics knowing last month x…

The analytics is similar to happy customers if you separate the customers from the bots.
Post reply on HN