Earlier quoted context omitted.
A lot of people justifiably care because making that information is their livelihood. Your entitlement to the labor of others is gross.
Ideas are copied by reading or hearing them. You can't own your ideas now, unless by own you mean horde. The perpetual creators rights you want extended are already artificial and require a non-trivial amount of our GDP to enforce and they still stifle future creation in a lot of areas. Most people are paid for doing things every day, they don't get to create one thing and never work again. Expanding creators compens…
Are large language models a threat to digital public goods?
131–140 of 163 posts
Re: Are large language models a threat to digital public goods?
#132Earlier quoted context omitted.
Or people who don't make anything. It's very easy to be generous with other people's work.
Techno communism essentially. "From each according to his ability, to each according to his needs". And just like that type of economic setup a handful of bros will reap the benefits - in this case sam altman's clique and others like them.
Re: Are large language models a threat to digital public goods?
#133Earlier quoted context omitted.
Perhaps there will be a drop in high value information in the public domain, but right now, I can't exactly see LLMs impacting the incentives for creation and sharing of that information. I don't see how LLMs would make someone go "oh well, AI is here, I might as well stop providing people with no-strings-attached high quality information", if the existence of search engines didn't make them stop already.
For years people have been making travel blogs based on where they've visited and the practical information they've discovered, like experiences of visiting attractions or good places to stay in cities or how they got from one place to another. They monetised with ads and affiliate links so they could travel more based on that income. In LLM land, they get no monetisation any more because nobody visits their sites, i…
I consider this to be a problem on its own, but it's not relevant here because:
> In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found.
That can't possibly be true, because if it were, there wouldn't be any travel blogs anymore today. All that travel spam has been made redundant approximately around the time Flickr was created, and every interesting location ever has been photographed from every interesting angle in a way neither me, nor you, nor your favorite travel blogger could ever hope to match. All the information they post has also been posted many times over by travel bloggers that came before.
The point being: travel information and photography is worthless commodity these days. Travel bloggers (or Instagrammers, or whatever) are not in the business of selling information. They're selling dreams and personal experiences. The photos and information are necessary as delivery vector ("social object"), but by themselves are worthless and not the point. The point is entertainment, and ideally getting you trapped in a parasocial relationship with the travel blogger/grammer, which gives them a recurring revenue stream.
> In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found.
It's the same model as with most other ad-monetized social media publishing. People will keep visiting them for the same reason they visit them now, and for the same reason they have their favorite youtubers and tiktokers. LLMs and other generative models don't change anything here, at least not short-to-mid-term, because they can't convincingly replicate human connection and keep it up for long.
(Also, I personally don't buy that travel instagrammers can actually sustain their travel lifestyle through ads and affiliate marketing and sponsorship deals. I suspect most are funded in some way, whether by family wealth or by services performed while traveling around.)
> The search engines actively supported these authors, by sending them people who needed the answers they had. (...) A LOT of the useful information on the web was built on similar feedback loops and they go away in LLM land.
Hard disagree. The only feedback loop this created in practice is the one that displaces quality information from the Internet - the combination of SEO and ad-based monetization means the most scummy players are the ones with most money to stalk every conceivable search query. The results speak for themselves: making a Google query for pretty much any topic of interest to general population will give you only content marketing sites - results that carry negative knowledge, as in if you waste your time reading them, you'll come more misinformed about the topic than you were before. If LLMs make all that go away, I'm 100% for it.
As for "A LOT of the useful information" - nope, can't think of a single case where ad/affiliate-supported site was a good information source, vs. just displacing a better free source.
Re: Are large language models a threat to digital public goods?
#134Earlier quoted context omitted.
Skill generating ideas is over rated. Implementation is when skill matters. I can imagine post scarcity Star Trek like ideas but not implement them. My skills at CE and SWE allow me to implement reasonable solutions there. Chasing ideas for the sake of chasing ideas is a form of bike shedding and premature optimization.
You can only say this right now where there are an abundance of ideas and less skill to implement them. But even implementation requires creative thinking and the ability to generate ideas. Now we want to take away the need for the first part. This is NOT like going from handwriting to printing or radio to TV. Something new is going on. I use ChatGPT for work in a limited way. If you haven’t tried it, I suggest that…
Re: Are large language models a threat to digital public goods?
#135Earlier quoted context omitted.
Then perhaps they should find other means of livelihood, instead of preventing the rest of the world from making full use of the information and technology available to it.
They will and less information will be put into places that are freely accessible. If it's put anywhere at all it'll be put behind login only/paywalled/unscrapable places that LLM's can't access.
The way I see it, there are roughly three groups of information providers:
1. Those who do it pro bono - because they feel like its a worthwhile thing to do, or because they believe in by "pay it forward", or otherwise because they haven't even thought that what they share is worth trying to extract rent from.
2. Those who do it "for free", as a way to lure people to where they can expose them to ads, affiliate marketing, upsells, or other such schemes - making money by being predators using information as bait.
3. Those who just put up a paywall, being up front that they're selling information, not giving it away.
(There's also a weird "in-between" group of publishers that are almost like 1., except they're being funded out of marketing budgets of companies that figure providing quality information is good advertising.)
LLMs don't change anything for group #1. They may compete with group #3, but that's business as usual, not anything transformative. The group that's directly affected is #2, which also happens to be the group that produces all the garbage on the Internet, so I'm actually very happy to see them forced to find a more useful way of making money. Since group #2 produces "information" that's arguably negative knowledge on the net, it's likely to improve the amount of quality information you'll be able to find on-line.
Re: Are large language models a threat to digital public goods?
#136Earlier quoted context omitted.
> They used stack overflow to make their case, and report that user engagement has gone down after the release of ChatGPT. Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be. So I look at the paper to see if they accounted for that, and find this. The moderation team and community in general on Stack…
I have anecdotal stories about aa friend corroborating this. He had been rebuked and experienced an unwelcoming response on Stack Overflow in response to his questions while trying to learn web development, and found ChatGPT a great teacher in comparison.
Re: Are large language models a threat to digital public goods?
#137Earlier quoted context omitted.
Ideas are copied by reading or hearing them. You can't own your ideas now, unless by own you mean horde. The perpetual creators rights you want extended are already artificial and require a non-trivial amount of our GDP to enforce and they still stifle future creation in a lot of areas. Most people are paid for doing things every day, they don't get to create one thing and never work again. Expanding creators compens…
I think that is somewhat off topic. I don't see why rent seeking via IP laws should be bad while doing it via provision of AI wouldn't be.
(BTW. that you can even make a system this way is a huge breakthrough that's not being talked about enough.)
But even if they were a mere database indexing copies of other peoples' IP, then - copyright issues notwithstanding - the de-bullshittifying of information retrieval process alone would be service worth paying a lot of money for.
Re: Are large language models a threat to digital public goods?
#138Earlier quoted context omitted.
It isn’t clear to me (other than it’s an open engineering problem) why LLs couldn’t also include attribution as part of training. Also tracking attribution could lead to some insights on how its internal representations in vector space are created.
That is true I see no reason obvious reason why the LL companies take pride in not being able to document ideation process. I have no justification but I feel it is deceitful not technical reasoning.
As an intuition pump, when I write "2+2 = " and you mentally complete it with "4", should I chastise you for not completing it with "4, as per ${your elementary class math textbook} and ${that other book you read as a kid}, corroborated by ${your first math teacher} and ${your parent} quoting ${some other work}"?
Re: Are large language models a threat to digital public goods?
#139Earlier quoted context omitted.
I have really enjoyed using https://phind.com , which includes attribution in its responses.
Phind’s base model, which is GPT3.5/4 doesn’t itself do attribution, it’s made to do that with prompt engineering which provides the most relevant materials on the web based on a word embedding vector search, and then asks it to reference each source in the answer.
I.e. in case of both the student and an LLM, correct citation doesn't actually mean the idea originates from the cited work - only that the work contains this idea.
Re: Are large language models a threat to digital public goods?
#140Earlier quoted context omitted.
What does your example have to do with this situation? The people in the bread supply chain get paid, the author of content we're discussing will never get paid by anyone, will never even get a bit if personal satisfaction from their analytics knowing last month x thousand people read that page and it hopefully helped them. It's completely zero reward, even worse it's completely zero feedback of any kind! This really…
> What does your example have to do with this situation? It's addressing GP's complaint about me pointing out the indirect and transactional nature of the interaction between information producer and consumer. > the author of content we're discussing will never get paid by anyone That's... not my problem? Bear with me here. > will never even get a bit if personal satisfaction from their analytics knowing last month x…