Live data from Hacker News

The false positive rate of AI detectors and its effect on freelance writers

authory.com

141–150 of 171 posts

Re: The false positive rate of AI detectors and its effect on freelance writers

#141

Earlier quoted context omitted.

Doesn't this mean Google has an AI-detector that is somehow accurately finding these? Otherwise the majority of its index would be falsely flagged. What do they use?

Microsoft is probably the best placed company to build a good detection method, having both a giant corpus of text predating LLMs from their Bing crawler, and a giant corpus of AI-generated text thanks to their relationship to OpenAI. But Google is probably a close second, with an even bigger archive of pre-LLM text (including the world's largest collection of digitized books), and supposedly a good LLM and lots of s…

It is impossible to create detector for LLM generated writing. The LLM detector technology must beat the LLM at its own game and if it can we can easily create a better LLM. Anyone who claims they can do it is sorely mistaken or flat out lying.

Re: The false positive rate of AI detectors and its effect on freelance writers

#143
post #46

There's a key sentence hidden right at the end of this article: > They were terrified of Google downranking any AI-generated content, and because of the apparently clear results from their AI checker, they couldn’t be convinced otherwise. It sounds to me like the client might even have believed the evidence that he provided them showing that he didn't use AI to write the work... ... but they're in the SEO game, and t…

This pattern has precedent. I scan releases with a bunch of Anti-Virus software. Not to find any virus, of course, but if one of the scanning softwares "detects" a virus, some customer is also going to "detect" the virus, and we have a problem.

Is it becoming less of a problem nowadays?

Windows has some good enough built in stuff, right? I can’t imagine using a third party scanner.

Re: The false positive rate of AI detectors and its effect on freelance writers

#144

Earlier quoted context omitted.

Microsoft is probably the best placed company to build a good detection method, having both a giant corpus of text predating LLMs from their Bing crawler, and a giant corpus of AI-generated text thanks to their relationship to OpenAI. But Google is probably a close second, with an even bigger archive of pre-LLM text (including the world's largest collection of digitized books), and supposedly a good LLM and lots of s…

It is impossible to create detector for LLM generated writing. The LLM detector technology must beat the LLM at its own game and if it can we can easily create a better LLM. Anyone who claims they can do it is sorely mistaken or flat out lying.

By the same logic I guess you could create one for previous generations of LLMs. "Looks like you need to update buddy...busted!" And across multiple industries a decent percentage of users will be using outdated versions...

Re: The false positive rate of AI detectors and its effect on freelance writers

#145
post #46

There's a key sentence hidden right at the end of this article: > They were terrified of Google downranking any AI-generated content, and because of the apparently clear results from their AI checker, they couldn’t be convinced otherwise. It sounds to me like the client might even have believed the evidence that he provided them showing that he didn't use AI to write the work... ... but they're in the SEO game, and t…

[deleted]

Re: The false positive rate of AI detectors and its effect on freelance writers

#147

Earlier quoted context omitted.

Microsoft is probably the best placed company to build a good detection method, having both a giant corpus of text predating LLMs from their Bing crawler, and a giant corpus of AI-generated text thanks to their relationship to OpenAI. But Google is probably a close second, with an even bigger archive of pre-LLM text (including the world's largest collection of digitized books), and supposedly a good LLM and lots of s…

It is impossible to create detector for LLM generated writing. The LLM detector technology must beat the LLM at its own game and if it can we can easily create a better LLM. Anyone who claims they can do it is sorely mistaken or flat out lying.

Your LLM detection method doesn't have to have a better understanding of human language than the LLM you're detecting, it just has to be able to identify patterns of speech typical of a particular LLM. Of course you have to repeat that work for each network and each significant update and finetune, but right now most people just use one of two ChatGPT versions.

And of course you don't tell people how your LLM detector works, don't provide a public API, and preferably leave it ambiguous whether it even exists. That solves most issues with LLMs using your detection method to get better in future versions.

That's counter to the business model of these LLM detection websites, but Google doesn't have a problem with keeping any potentially existing models inhouse and locked down.

Re: The false positive rate of AI detectors and its effect on freelance writers

#148
I'm having to start having doubts, recently Riot games had a controversy where one of its artists was caught using AI for art, just like with cheating in pro gaming/sports a lot of people say no but then turn out to really be doing it.

Re: The false positive rate of AI detectors and its effect on freelance writers

#149
post #46

There's a key sentence hidden right at the end of this article: > They were terrified of Google downranking any AI-generated content, and because of the apparently clear results from their AI checker, they couldn’t be convinced otherwise. It sounds to me like the client might even have believed the evidence that he provided them showing that he didn't use AI to write the work... ... but they're in the SEO game, and t…

This pattern has precedent. I scan releases with a bunch of Anti-Virus software. Not to find any virus, of course, but if one of the scanning softwares "detects" a virus, some customer is also going to "detect" the virus, and we have a problem.

Microsoft Defender "detects" a "virus" on any .exe or .dll of unknown provenance.

The solution is code signing. Pay whatever you need to pay for a certificate and sign all your releases.

Re: The false positive rate of AI detectors and its effect on freelance writers

#150

Earlier quoted context omitted.

This pattern has precedent. I scan releases with a bunch of Anti-Virus software. Not to find any virus, of course, but if one of the scanning softwares "detects" a virus, some customer is also going to "detect" the virus, and we have a problem.

Is it becoming less of a problem nowadays? Windows has some good enough built in stuff, right? I can’t imagine using a third party scanner.

People still have anti-virus software though. How good window's defender is is not relevant. What matters is the prevalence of various anti-virus software for your target userbase. Pretty sure various common software install tools give options for installing things like McAfee or Avast. Or they come preinstalled on store-bought systems.
Post reply on HN