Live data from Hacker News

Generative AI could make search harder to trust

wired.com

241–250 of 255 posts

Re: Generative AI could make search harder to trust

#241
post #237

Earlier quoted context omitted.

It's not that simple. Originally OpenAI released a model to try and detect whether some content was generated by an LLM or not. They later dropped the service as it wasn't accurate. Today's models are so good at text generation it's not possible in most cases to differentiate between a human and machine generated text.

Well they could just not allow prompts that seem to participate in blogspam. If they wanted to stop it they definitely could. Their argument is that since it's centralized, things like that are possible (while with llama2 you can't), they do "patch" things all the time. But since blobspam are contributing to paying back the billions microsoft expects they're not going to.

> Well they could just not allow prompts that seem to participate in blogspam.

Unfortunately, any question that lots of people legitimately ask will also be a prompt for blogspam.

Re: Generative AI could make search harder to trust

#242

Earlier quoted context omitted.

Yes. Only the enlightened like me should get it…smh.

We don't hand dangerous tools to people without proper training. I don't see why this would be any different.

Really? Anyone can go into a Home Depot and get many dangerous tools with no training.

Re: Generative AI could make search harder to trust

#243

Earlier quoted context omitted.

Because everyone has at least tried to lie - not necessarily for malicious reasons, e.g. white lies - a few times in their lives and know it’s not easy coming up with plausible sounding lies. It takes effort. Sure it might be effortless for some psychopaths but there are a limited number of such people. Overall the internet was surprisingly usable despite all the liars - even now with SEO and content farms; I think e…

Again... > Because everyone has at least tried to lie - not necessarily for malicious reasons, e.g. white lies - a few times in their lives and know it’s not easy coming up with plausible sounding lies. How do you know that?

> How do you know that?

Common knowledge.

Regardless, the amount of false information on the internet used to be limited by the amount of people producing it. Some of it is deliberately produced misinformation. Some of it are just mistakes due to carelessness, others due to willful negligence - example of the latter would be content farms who make their money off flooding search engine result pages and pushing ads in your face when you visit their website; who couldn’t care less what if you got what you came for.

Despite all that the internet was still useful. There are enough people posting accurate information such that the signal to noise ratio is good enough.

LLMs will change all that. They have the potential to flood the internet with hallucinated falsehood. And because bad LLMs are cheaper and easy to create and operate, those will be the majority of LLMs used by aforementioned content farms.

Re: Generative AI could make search harder to trust

#244

Earlier quoted context omitted.

Sadly Reddit is sort of monetized, as we have people selling accounts for what I believe are spamming and propaganda purposes.

Yes, but it's reddit monetizing, not the users. The users can be astroturfing, but there's no point astroturfing some types of content like game guides I hope.

Some users monetize it by selling their karma rich accounts to those astrosurfers you mentioned - or to spammers.

They currently manufacture these karma rich accounts by reposting popular posts and comments. LLMs will soon be (or already are) another way to karma farm.

Re: Generative AI could make search harder to trust

#245

Earlier quoted context omitted.

Not saying they should, but if they wanted to they could have an API that allows you to check whether some text was generated by them or not. Then Google would be able to check search results and downrank.

It's not that simple. Originally OpenAI released a model to try and detect whether some content was generated by an LLM or not. They later dropped the service as it wasn't accurate. Today's models are so good at text generation it's not possible in most cases to differentiate between a human and machine generated text.

I mean they would literally just store hashes of all the content that they generate and compare it vs their database. No AI involved.

Re: Generative AI could make search harder to trust

#246
post #168

Earlier quoted context omitted.

Honestly, I would love to see a LLM based plugin that would take a website, remove all the tracking garbage and filler, and just give me back a no-frills static html+CSS site that looks like it was made in 1995. "Oh, you want a guide to writing your own loss function for Tensorflow? Here's an FAQ that could have existed on comp.lang.python3.tensorflow"

So you want a... reader mode plugin? How would it use an LLM beneficially?

Does reader mode eliminate "filler"?

Re: Generative AI could make search harder to trust

#247

Earlier quoted context omitted.

It's not that simple. Originally OpenAI released a model to try and detect whether some content was generated by an LLM or not. They later dropped the service as it wasn't accurate. Today's models are so good at text generation it's not possible in most cases to differentiate between a human and machine generated text.

I mean they would literally just store hashes of all the content that they generate and compare it vs their database. No AI involved.

If it's a straightforward hash, that's easy to evade by modifying the output slightly (even programmatically).

If it's a perceptual hash, that's easy to _exploit_: just ask the AI to repeat something back to you, and it typically does so with few errors. Now you can mark anything you like as "AI-generated". (Compare to Schneier's take on SmartWater at https://www.schneier.com/blog/archives/2008/03/the_security_...).

Re: Generative AI could make search harder to trust

#248

Earlier quoted context omitted.

AI has so completely disrupting Search that it’s destroyed leading platforms effectiveness in a matter of months. But because of its current lack of optimization for accuracy, we shouldn’t consider it disruptive because it’s not yet proven technology? You can call it dangerous but you can’t call it useless. It’s also only going towards improvement from here, including drastic reductions in hallucinations. You have to…

> It’s also only going towards improvement from here Why?

Marx, Kondratiev, Schumpeter, creative destruction, 6th wave, blah blah blah

Re: Generative AI could make search harder to trust

#249

Earlier quoted context omitted.

Not saying they should, but if they wanted to they could have an API that allows you to check whether some text was generated by them or not. Then Google would be able to check search results and downrank.

It's not that simple. Originally OpenAI released a model to try and detect whether some content was generated by an LLM or not. They later dropped the service as it wasn't accurate. Today's models are so good at text generation it's not possible in most cases to differentiate between a human and machine generated text.

It would be easy to workaround using other open source models. You use GPT-4 to generate content and then LLAMA-2 or sth else to change the style slightly.

Also, it would require OpenAI to store the history of everything that it's API has produced. That would be in contrast with their privacy policy and privacy protections.

Re: Generative AI could make search harder to trust

#250

Leaving aside the article to discuss the source for a moment. When did Wired become so antitech? There are good critical viewpoints but most of the articles they are putting out at this point read like bitter diatribes. Which is a shame because they used to be an excellent publication.

I saw a comment describing this increasing luddite response to new tech developments and it resonated with me. Essentially, the comment made the point that tech is advancing so fast these days that most people are unable to keep up with the pace of these radical changes. And the natural reaction to that for many is to reject these "advancements" or at least look upon them with cynicism and skepticism.

You don't think it has anything to do with the fact that 'advancements' these days come with mandatory accounts/subscriptions, egregious data harvesting, and user-hostility to keep people in line with what companies want rather than what people want?
Post reply on HN