Live data from Hacker News

The AI bullshit singularity

successfulsoftware.net

71–80 of 187 posts

Re: The AI bullshit singularity

#71

Succinctly stated and something that resonates strongly with me. In the last internet revolution (web search), results started high quality because the inputs were high quality - bloggers and others just wanted to document and share knowledge. But over time, many interests (largely commercial) figured out how to game the system with SEO, and quality of search results has decreased as search's incentive structure led…

I guess some of it might depend on how good the AI-generated content gets, and also how good the AI gets at detecting AI-generated content.

If the AI was good at detecting it, it wouldn't matter if the AI-generated content sucked, yes?

Even a low probability of detection would help. Let's say our algorithm is 50% likely to detect AI junk. That means that half the junk data won't make it in to model. Even 20% would probably be worthwhile, especially if it also threw away human-generated junk (and let's be realistic here: there is, and always has been, no shortage of terrible and/or wrong human-generated content).

Let's say those crappy filler paragraphs that get stuck between pictures on meme clickbait pages... I suspect most of those people have already been replaced, but that prose was horrible long before the current LLM boom.

I suspect there is a lot of effort being expended right now on ways to ensure the training data (whether human or AI generated) isn't shite.

Re: The AI bullshit singularity

#72

Succinctly stated and something that resonates strongly with me. In the last internet revolution (web search), results started high quality because the inputs were high quality - bloggers and others just wanted to document and share knowledge. But over time, many interests (largely commercial) figured out how to game the system with SEO, and quality of search results has decreased as search's incentive structure led…

Unlike the low background radiation steel, high quality content will still continue to be created. Arguably even at a faster rate with these powerful tools. The proportion will drop ofcourse but that only means curation will become king. With search this curation was difficult, because the value was so low you could only do it profitably if it was completely automatic, and that was difficult.

An optimistic take would be that LLMs will make curation so much more valuable, that it will be done much better. If the wider world gets to use this curation to limit spam, and highlight good work, it would be amazing for the world.

Re: The AI bullshit singularity

#73
post #29
post #23

Earlier quoted context omitted.

Cryptography can make content you create marked. But how would you use it to mark content other (non cooperative) people make?

Ya anyone that tells you crypto solves the problems of bad actors is delusional or pawning a bill of goods.

No, you can filter bullshit by trusted public keys. This isn’t a new problem. It just now needs to go mainstream.

Re: The AI bullshit singularity

#74
post #18

Earlier quoted context omitted.

This is such a wrong take. Even if the data is 100% synthetic, you can still hill climb to new mountains. If you don't believe me, look at evolution. It doesn't matter if we no longer have 100% human art as input. This is the worst these systems will ever look and feel, and they're only going to improve. I'd be willing to do a longbets on this one.

Evolution has a clearly defined fitness function (number of offspring that are produced). What is the fitness function for generated art?

Exactly, the fitness function is "make a picture which looks like these other pictures", which is used as a proxy for "make a picture that a human would find appealing by making it like all these pictures made by humans". The former definiton will always hold true, but the latter actually desired definition breaks down once your training set is contaminated with imitations of imitations of imitations.

Re: The AI bullshit singularity

#75

Earlier quoted context omitted.

Good content won't disappear necessarily, it could however be drowned out by BS as the author states. Its the needle in a haystack problem, where AI is used to increase the size of the haystack, making the needle (i.e. quality content) harder to find.

The Internet is a pull model. I can go to Gwern's website directly and not care that most other websites have crap on them. People choose to use push models for content through meta properties, tiktok, and aggregators like reddit and HN, but nothing is forcing them to. If they push enough bad content, people won't keep using them. Already happened with Facebook and Reddit predecessors, probably happening to Reddit no…

Perhaps LLMs will figure out the same thing.

Re: The AI bullshit singularity

#76
post #18

Earlier quoted context omitted.

This is such a wrong take. Even if the data is 100% synthetic, you can still hill climb to new mountains. If you don't believe me, look at evolution. It doesn't matter if we no longer have 100% human art as input. This is the worst these systems will ever look and feel, and they're only going to improve. I'd be willing to do a longbets on this one.

Evolution has a clearly defined fitness function (number of offspring that are produced). What is the fitness function for generated art?

Number of works that are released to the internet! At the end of the day, there is a human who has an idea in mind and is using an LLM to realize it. They are tweaking their prompt until they achieve their vision, then publishing only those successes. The resulting art can then be used with the prompt as training data.

Re: The AI bullshit singularity

#77
The web already contains vastly more information than you could ever even begin to read or look at in a lifetime. Most of it is total garbage. You have to use a trust filter to find things worth looking at.

You're arguing that there isn't a trust filter that can do that, but there will be, and it will probably be an AI.

Re: The AI bullshit singularity

#78
post #18
post #5

The same goes for image generation models, AI art already has a tendency to veer into the same clichés and those are only going to get reinforced if newer models are trained on newer scrapes which now include the million hyper-derivative AI images being uploaded to places like DeviantArt, Twitter and Pixiv every day. Those vendors who got in early have a moat in the form of untainted scrapes, but they'll eventually n…

This is such a wrong take. Even if the data is 100% synthetic, you can still hill climb to new mountains. If you don't believe me, look at evolution. It doesn't matter if we no longer have 100% human art as input. This is the worst these systems will ever look and feel, and they're only going to improve. I'd be willing to do a longbets on this one.

It's getting really tiresome to see tech bros talk about biology as if they know what they're talking about.

Re: The AI bullshit singularity

#79
The Dead Internet Theory was only slightly ahead of its time.

Used to be real people pretended to be girls on the internet. These days I can’t even get an honest real fake person pretending to be an attractive female on LinkedIn.

Re: The AI bullshit singularity

#80

Succinctly stated and something that resonates strongly with me. In the last internet revolution (web search), results started high quality because the inputs were high quality - bloggers and others just wanted to document and share knowledge. But over time, many interests (largely commercial) figured out how to game the system with SEO, and quality of search results has decreased as search's incentive structure led…

I think it’s clear that LLMs cannot be the end state of this technology, and we will need systems that can reason and develop hypotheses and test them internally. These systems may benefit from more curated datasets (such as those collected before the bullshit wave began) along with real world interaction data from YouTube and robotics. Such systems could eventually be used to rank web pages for their bullshit level,…

> It’s just really clear that a giant text averaging machine can only go so far

It's not really a text averaging machine, it's a pattern matching machine.

Right now the "depth" of the patterns it can match can only go so far, but in a few years with more advances in chips and memory the depth is going to increase and the patterns it can match will fan out accordingly.

Post reply on HN