Live data from Hacker News

Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

lajili.com

261–270 of 392 posts

Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

#261
post #223

Earlier quoted context omitted.

I don't have much knowledge around AI, but from what I can tell, it's dependent on inputs from across the web right? If so, then as the use of ChatGPT grows, it'll slowly get consumed back into the model. With enough iterations, it will start to veer more and more away from recognizable human speech/thought, like a recursive game of telephone. Unless I'm completely misunderstanding how ML works, which very well may b…

ChatGPT itself is being trained on curated content, it is clearly not trained on unscreened internet sites - this is easy enough to establish if you ask it questions around hot topic issues in the conspiracy groups - it gets the correct mainstream answer. I suspect we will see the rise of both groups of machines, curated A.I.s and A.I.s just trained on anything, which should be entertaining.

Curated content by non native Enlish speakers earning as little as $2.00 an hour. What can possibly go wrong?

Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

#262

> The best I can hope for is that some hacker collective manages to make an open source version of it, that can be more trustworthy than the current one. It's always interesting to me when someone makes an assert like this for a Big Data technology. If this were tractable, we'd have an open-source Google alternative right now that someone would have built for the sheer joy of being the folks that took on Google. But…

Unlike Google, you can download Wikipedia and use it offline. I hope to see the same thing happening for a useful LLM. Of course that’s not feasible right now, but hopefully those costs will come down and we do actually get to run a self-hosted version of this.

Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

#263

Earlier quoted context omitted.

The worst part about this is that if there is another set of bots that tries to generate engagement, then the training data isn't coming from humans either. You have one set of actors spamming. And another set of actors upvoting stuff, predominantly their own but maybe also other random posts. So the resulting posts don't necessarily even cater to humans. It will be real online hellscape.

The web will transition strongly to verified identities, like we have with SSL certs. Along with filtering out people who use AI to post under a verified identity and get caught, It’s the only way to help ensure you’re reading actual human content.

If the web degenerates to the point where verified identities are required, then it really and truly will have died.

Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

#264

Earlier quoted context omitted.

The difference is we can improve the AI to be more accurate, and I suspect before long it’ll generate better content than a human would that’s verifiable with citations. There may come a time where writing is done by a machine much as a calculator does our math. But knowledge maybe shouldn’t be canonically encoded in ascii blobs randomly strewn over the web - maybe instead of accumulated knowledge needs to be structu…

The model needs known "good" feedback to improve. The problem is that the quality of its training data declines with the more output produced. It's rather inevitable that we'll be drowning in AI generated garbage before long. A lot of people are confusing LLMs with true intelligence.

That’s why I think knowledge needs to be better structured than blobs of text scattered everywhere. An AI can be more than an LLM, Wolfram posted recently about that. You can use the LLM to convert a question into a semantic query and a semantic validator and check and amend and provide a semantic knowledge graph explaining an answer and the LLM can convert it back to meat language. I think people confuse LLM with true intelligence, but the cynics also confuse LLM with a complete and fixed point solution.

Your point also seems to assume no curation can happen on what is ingested. Simply because that might be what’s happening now you could also simply train the LLM on known good sources and be as permissive or restrictive as is necessary. Depending on how good the classifiers are for detecting LLM output (openai released on recently) or other generated / automatically derived content you can start to be more permissive.

My point is people seem to be blinded by what is vs what may be. This is not the end of the development cycle of the tech, it’s the pre-alpha release by the first meaningful market entrant. I’d be slower to judge what the future looks like rather than assuming everything stays fixed in time as it is.

Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

#265

Earlier quoted context omitted.

Maybe information retrieval against unstructured text isn’t the right model? Maybe it’s time for google search to die.

AI generated content almost certainly will kill it as we know it. I don't expect the interface to change, but I expect Google's AI will "decide" what gets placed in search results, and where.

I think search itself was always a hack for how to ask human knowledge a question. It’s a great way to find a specific page of documentation. That’s really not what people want to do 99% or the time they use google.

Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

#266

> The best I can hope for is that some hacker collective manages to make an open source version of it, that can be more trustworthy than the current one. It's always interesting to me when someone makes an assert like this for a Big Data technology. If this were tractable, we'd have an open-source Google alternative right now that someone would have built for the sheer joy of being the folks that took on Google. But…

Unlike Google, you can download Wikipedia and use it offline. I hope to see the same thing happening for a useful LLM. Of course that’s not feasible right now, but hopefully those costs will come down and we do actually get to run a self-hosted version of this.

The tricky thing about data is that the world constantly changes. A downloaded Wikipedia has a lot of value, but it does grow stale. And it has the advantage of being a repository of relatively static facts in a way that, say, a search engine is not.

Search engines (and I suspect a ChatGPT-style engine, if one wants to talk about it about current events, things currently available, or other topics of the day) have to be continuously refreshed to be relevant. So many things that those engines are used for frequently (including the keyword "ChatGPT" itself) had no definition months ago, let alone an inaccurate definition.

Most data isn't static like code; it must be continuously re-invested in to stay relevant.

Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

#267
post #145
post #136

Each webpage should express a piece of metadata that is content-quality. * primary - The content on this page is a primary source * secondary - The content on this page is high quality, but which is primarily research based and may thus be tainted by unknown sources * bot - The content on this page is mostly automatically generated by one or more ML models with some human curated improvements HTML elements can also h…

If I operate a webpage, and I incorporate this content quality metadata tag, why would I put anything other than "highest possible quality"?

[deleted]

Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

#268
post #58

Imagine a world where the only content you see is from publishers that you trust, and that your friends trust, and their friends, to maybe 4 or 5 hops or so, and the feed was weighted by how much they are trusted by your particular social graph. If you start seeing spammy content, you downvote it, and your trust level from that part of your social graph drops, and they are less likely to be able to publish things tha…

This is similar to what post.news is doing.

Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

#269
post #202

Earlier quoted context omitted.

Worse, perhaps. Kessler Syndrome will eventually resolve itself as junk falls out of orbit over time, or new methods for cleaning it up are developed. Information, once buried in noise, becomes unrecoverable without a source of known truth for correlation.

Curation is the answer. The more junk, the better off are the one curation for quality, relevancy and human interest.

Curation that tracks the provenance. If we receive a string of text by itself we can't do much about it. We need to know from where it came from, whether it was written by a human, etc

Re: Chat GPT is the birth of the real Web 3.0, and it's not going to be fun

#270

An endless loop of AI generated content that gets posted to the web as original human generated content, with LLMs getting re-trained on this content and spitting out more content that also gets re-posted, resulting in a cesspool of BS masquerading as organic knowledge. I'm old enough to remember when Google provided meaningful search results rather than just SEO spam, the problem is about to get an order of magnitud…

> An endless loop of AI generated content that gets posted to the web as original human generated content, with LLMs getting re-trained on this content

Man I'm so tired of this very obvious observation. I wouldn't think a company smart enough to create an AI would also be dumb enough to fall into a pitfall that even the most casual observer can identify.

Post reply on HN