Live data from Hacker News

Detecting LLM-Generated Texts with “Classical” Machine Learning

blog.lyc8503.net

121–130 of 184 posts

Re: Detecting LLM-Generated Texts with “Classical” Machine Learning

#121
> Sounds promising, right? I spent some time trying [perplexity], but results were disappointing—plenty of false positives and false negatives, and no reasonable threshold could be set.

Perplexity was widely considered SOTA in 2022. One part of it is because everyone was evaluating on open models or closed models that were still close (i.e. GPT-2 vs. GPT-3.5). Today, the gap is so much wider between the models you can use to compute perplexity and the frontier models people actually use.

Also so many AI text detection papers used a strawman RoBERTa baseline that was very undertrained for the task.

The synthetic mirrors method for data generation used here is the same as what we use at Pangram. Good blog post, thank you for sharing!

Re: Detecting LLM-Generated Texts with “Classical” Machine Learning

#123

Earlier quoted context omitted.

Wrong. Effective sampling (I.e high temperature like temp 10) with the corresponding sampler stack that enables this coherently destroys all attempts to detect it. There are many more ways like this involving manipulating the logprobs

I’m not sure what you’re saying I’m wrong about. The comment I was responding to, about changing LLM output, referred to prompting, not temp/sampling tricks. I’m not aware of Pangram being beat by clever prompting. There’s some interesting work on creative writing using contrastive prompt techniques, but I haven’t seen it tried as evasion. Even if you control temp and sampling, they’re not magic. If you raise the tem…

As min-p approaches 1, the temperature you can get away with approaches infinity. Also more modern samplers like top-n-sigma are explicitly designed to get away with temperature of infinity.

Re: Detecting LLM-Generated Texts with “Classical” Machine Learning

#124
post #117

Earlier quoted context omitted.

IIRC, the big names in LLMs have no real interest in cloaking the LLM-nature of the text, Google adds deliberate watermarks to text, OpenAI developed a watermark for text but reportedly arent't actually using it.

Considering [1], I’m going to challenge that their techniques are currently even mildly effective. Given the absolute academic malpractice these papers are pushing, I’m calling BS; while they want to watermark it, they clearly aren’t actually able to. For images. Which are drastically easier than text. Their interest is irrelevant in the face of technical impossibility. And that’s before you get into other people who…

> Their interest is irrelevant in the face of technical impossibility.

I'm responding to "If the data was separable in this way, you would equally be able to train an AI to mask those signs.": yes, if you wanted to you could, the big names clearly don't consider masking to be a priority.

> But it’s absolutely unusable for something like “did someone cheat”.

This is the one case where I'd most expect it to succeed:

I suspect most of the people who do want to cloak-to-cheat, don't have the skills to do so; I also suspect most of them are so unaware of what they don't know that they won't even ask an LLM to write cloaking software for them.

Re: Detecting LLM-Generated Texts with “Classical” Machine Learning

#125
post #8

Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Images, absolutely, there are…

So you’re saying that the linked article’s findings are implausible? Is the article fake, then, in your opinion?

Re: Detecting LLM-Generated Texts with “Classical” Machine Learning

#126
post #98

Earlier quoted context omitted.

With sufficient information you can derive a signal even in the presence of overwhelming noise. Assuming the noise is not perfectly correlated with the signal this is always possible. Schemes like GPS, CDMA and DSSS are based upon this concept. GPS in particular is quite impressive in its ability to recover information that is received below the thermal noise floor.

There has to be a signal to detect it. Take this sentence: Bob went to the store to buy milk. Was that AI generated or not? There simply isn't a signal there. The problem isn't noise, the problem is, is there even a signal to begin with. Sure, you might be able to recognize the quirks of a specific LLM just as you recognize the quirks of a particular person, but as the number of LLMs proliferate, then the signal turn…

but.... the LLMs are actually all trained on approximately the same stuff, and tend to have similar quirks. In the human world, writers develop recognizable voices, which are detectable and classifiable (as in the article we have all supposedly read). Furthermore, we don't necessarily care about telling one LLM from another, just that they aren't human. That's different from trying to identify one human amongst a sea o fhumans, or one bot form within a sea of bots.

Re: Detecting LLM-Generated Texts with “Classical” Machine Learning

#127
post #8

Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Images, absolutely, there are…

If you have access to the detector, you can formulate a generative solution that avoids being flagged. Which gets me wondering why don’t model providers do that? There must be something about that that destroys semantic weights somehow.

Re: Detecting LLM-Generated Texts with “Classical” Machine Learning

#128
post #8

Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Images, absolutely, there are…

The article discusses a technique by which the author achieves high accuracy at detecting AI written text. Unless you have a problem with their experimental method, this is the opposite of tarot card reading.

> we are well into undetectable sophistication with today's models

The article directly contradicts this, as do you, in your previous paragraph: "Sure you might be able to detect today's tells". The article is literally about a technique that detects today's tells.

Your comment is mostly expressing doubt that this technique will work reliably in the future, but it's framed as opposition to the article, which it's not: the article is about detecting today's AI-written text, at which it seems to be quite successful.

Re: Detecting LLM-Generated Texts with “Classical” Machine Learning

#129
post #8

Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Images, absolutely, there are…

If you have access to the detector, you can formulate a generative solution that avoids being flagged. Which gets me wondering why don’t model providers do that? There must be something about that that destroys semantic weights somehow.

Why would sounding human be a goal rather than a byproduct of trying to communicate efficiently?

Re: Detecting LLM-Generated Texts with “Classical” Machine Learning

#130
post #64

Earlier quoted context omitted.

Whether a text was written by a human or not is just a single bit of information. So you can't rule out its detectability a priori, since even the shortest text contains more information than that. As long as LLMs are used to write texts humans wouldn't want to write if they could help it (that's why they're getting an LLM to do it, after all), they'll remain detectable. Even if the reasoning might end up equivalent…

That's like saying whether or not you're going to fall in love this year is just one bit of information, so you might be able to read it from astrology. Yeah, sure, it might happen for some people with a certain star sign. But across the population there is zero reason to believe that there is a) any significant correlation and b) enough data variation in to even distinguish classes of humans.

The 80% accuracy from the article would be one reason to believe there's significant correlation, no?
Post reply on HN