An endless loop of AI generated content that gets posted to the web as original human generated content, with LLMs getting re-trained on this content and spitting out more content that also gets re-posted, resulting in a cesspool of BS masquerading as organic knowledge. I'm old enough to remember when Google provided meaningful search results rather than just SEO spam, the problem is about to get an order of magnitud…
Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
231–240 of 392 posts
Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
#232It's always interesting to me when someone makes an assert like this for a Big Data technology.
If this were tractable, we'd have an open-source Google alternative right now that someone would have built for the sheer joy of being the folks that took on Google. But open source doesn't work that way because code is download-once, use-forever, but data is continuously changing and costs perpetual money to update and maintain. "Open source data" looks like Wikipedia, and the world won't sustain more than a few of those; Wikipedia has about 100,000 active editors.
So instead of some hacker-alternative-to-Google techno-utopia idea, we've got plenty of open-source crawlers and a handful of services paying the bills via rent-seeking their database and, often, advertising. No reason to think a ChatGPT-heavy future will be different.
Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
#233Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
#234Imagine a world where the only content you see is from publishers that you trust, and that your friends trust, and their friends, to maybe 4 or 5 hops or so, and the feed was weighted by how much they are trusted by your particular social graph. If you start seeing spammy content, you downvote it, and your trust level from that part of your social graph drops, and they are less likely to be able to publish things tha…
Trust is not transitive. My friends reshare all sort of crazy memes, and their friends are even worse. Just because you know someone doesn't mean they're good at reading the news or understanding what's going on in the world.
I have friends whose movie recommendations I trust but whose restaurant recommendations I don't, and vice versa. I have friend that I trust to be witty but not wise and others the opposite.
A system that tried to model trust would probably need to support tagging people with what kinds of things you trust them in.
Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
#235Earlier quoted context omitted.
One solution is pretty simple: pedigree. Divide up the 'net into trusted and untrusted sources. Make the trust ratings public. Use search tools and corpuses such as the Google Books dataset to source "knowledge" back to pre-Internet roots, when necessary. In short: bring academic reputation back and bring it back hard. It will make for a more elitist web, but given that even without ChatGPT we've had a problem with w…
That means anonymously posting on stackoverflow will be gone...
What it means is the bar for becoming a new StackOverflow contributor (or Reddit admin, or Wikipedian) might become much, much higher. "Oh, you want to contribute your first post? Show me the bicycles in this image, find the letters in this image, and provide the names of two existing Stack Overflow users with over 1000 karma who can vouch for you, and also you see a tortoise on its back, baking in the sun. You're not helping it. Why aren't you helping it?..."
Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
#236> The best I can hope for is that some hacker collective manages to make an open source version of it, that can be more trustworthy than the current one. It's always interesting to me when someone makes an assert like this for a Big Data technology. If this were tractable, we'd have an open-source Google alternative right now that someone would have built for the sheer joy of being the folks that took on Google. But…
Not just the data, but also hosting the service and keeping it available.
Open source works because the marginal cost of code is essentially zero. But the marginal cost of serving users is definitely not zero.
Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
#237Not a super well thought out article. Example: lots of speculative complaints that ChatGPT will lead to an explosion of low quality and biased editorial material, without a single mention of what that problem looks like today (hint: it was already a huge problem before ChatGPT). Ditto with the “ChatGPT gave me wrong info for a query” complaint. Well, how does that compare to traditional search? I’m willing to believe…
I see this counter-argument all the time and it makes no sense to me.
Yes, the web is already filled with SEO trash. How is that an argument that ChatGPT won't be bad? It's a force multiplier for garbage. The pre-existence of garbage does not at all invalidate the observation that producing more garbage more efficiently is even worse.
Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
#238This article lost me with it's definitions of "Web 1.0" and "Web 2.0," which are totally divorced from reality.
How so? I was curious if my memory was wrong but wikipedia seems to agree with me: https://en.wikipedia.org/wiki/Web_2.0 , what is your definition of web 2.0?
Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
#239Earlier quoted context omitted.
Okay so you build a knowledge graph on top of the internet archive. Now you are struggling to prioritize the resources necessary to capture long-tail content that doesn't mesh easily into popular corpuses. I imagine this would lead to the library equivalent of an echo chamber.
I was thinking more of a federated "webring" structure, with some content being present in more than one node, and where maintenance and curation are distributed (and gathered independently) among nodes. The nation of, say, Japan, has limited interest in funding an american noprofit today; but they would likely have a great deal of interest in funding an equivalent focused on Japanese content, for example.
So now you get into the issue of haves and have nots. Who is allowed to be considered an authorized archivist from a robots.txt perspective? Or what happens if an archivist becomes blacklisted for not respectfully crawling? How do national sanctions affect the Internet Archive of Russia? I imagine there would be a certification process and it would probably cost some money.
It's an interesting topic and I'm simply looking at the weak spots. I'm not against the overall concept though.
Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
#240An endless loop of AI generated content that gets posted to the web as original human generated content, with LLMs getting re-trained on this content and spitting out more content that also gets re-posted, resulting in a cesspool of BS masquerading as organic knowledge. I'm old enough to remember when Google provided meaningful search results rather than just SEO spam, the problem is about to get an order of magnitud…
In the back of my mind, I have a hope that it will lead to the collapse of the platform internet and a return to smaller trusted communities and boards.
> The dark forest theory of the web points to the increasingly life-like but life-less state of being online.Dark Forest Theory of the Internet by Yancey Strickler Most open and publicly available spaces on the web are overrun with bots, advertisers, trolls, data scrapers, clickbait, keyword-stuffing “content creators,” and algorithmically manipulated junk.
> It's like a dark forest that seems eerily devoid of human life – all the living creatures are hidden beneath the ground or up in trees. If they reveal themselves, they risk being attacked by automated predators.
> Humans who want to engage in informal, unoptimised, personal interactions have to hide in closed spaces like invite-only Slack channels, Discord groups, email newsletters, small-scale blogs, and digital gardens. Or make themselves illegible and algorithmically incoherent in public venues.