Live data from Hacker News

Scanning for Pangram Errors

veryfineprint.substack.com

11–20 of 44 posts

Re: Scanning for Pangram Errors

#11
Note that any good binary classifier (like "this is AI writing" for Pangram) should maximize the following four probabilities:

1. P(actually true | predicted true) (= true positive rate, recall, sensitivity)

2. P(actually false | predicted false) (= true negative rate, specificity)

3. P(predicted true | actually true) (= positive predictive value, precision)

4. P(predicted false | actually false) (= negative predictive value)

If you just try to minimize the share of false positives, you are only maximizing 1.

The measure which maximizes all four values is the binary Pearson correlation, also called "phi coefficient" or "MCC". It is 1 if the classifier is always correct, 0 if it guesses randomly (its predictions are statistically independent of the facts) and -1 if it always predicts the opposite of the true value.

Re: Scanning for Pangram Errors

#12
post #5

Interesting methodology. The results for Pangram are surprisingly good. These tools are definitely not 100% perfect, which is the primary complaint used to dismiss them. However the error rate is also getting impressively low. In most cases I see socially and online, there is a high suspicion that the content is AI generated before someone thinks to submit it to Pangram. It’s used on-demand as a tool to confirm suspi…

I think outside of a few niche uses, e.g. schools/universities, services like Pangram will struggle.

Currently people care because so much of AI output is still poor quality. But given I personally know of someone with published AI-generated works where the reviews are gushing and positive across the board, I doubt people will care whether works are AI generated for very long vs whether the text is good.

A way of certifying a work as above a certain literary quality would likely have far more longevity, as I doubt the market of people who care whether a given work is AI or not irrespective of quality will be particularly large.

Re: Scanning for Pangram Errors

#13
post #9

> Pangram boasts a false positive rate of 1 in a 10,000. That is, if Pangram says a block of text is AI there is only a one in ten thousand chance that it was written by a human. That'd be if they had a false discovery rate of 1/10,000. If for instance: * 100,000 samples are tested * 100 of which are AI-generated, the rest human-written * Pangram flags 50 of the AI-generated samples (true positives) * Pangram also fl…

Populare Science is built on a foundation of misunderstanding p-value.

https://www.graphpad.com/guides/prism/latest/statistics/comm...

(OP isn't exactly talkigh about p-value, because it's not comparing means, but the main mathematical structure of the Bayesian error is the same)

Of course, it's unclear if the article is using the wrong term, or the wrong definition, or if Pangram is just lying in the first place.

Re: Scanning for Pangram Errors

#14
Pangram is witchcraft to me. The way they can correctly detect AI writing from small samples, with so little statistical signal, is crazy. I've seen it correctly detect AI writing even when people use those "humanizer" rewriting skills that remove the hallmarks of AI writing and make the text indistinguishable from human text. But not to Pangram.

Re: Scanning for Pangram Errors

#15
post #9

> Pangram boasts a false positive rate of 1 in a 10,000. That is, if Pangram says a block of text is AI there is only a one in ten thousand chance that it was written by a human. That'd be if they had a false discovery rate of 1/10,000. If for instance: * 100,000 samples are tested * 100 of which are AI-generated, the rest human-written * Pangram flags 50 of the AI-generated samples (true positives) * Pangram also fl…

The classic. If there is one concept people should learn from stats it should be base rates and the Bayes update. It's very difficult to teach though. The mind just jumps to that interpretation, and even if you teach it, people just go right back to interpreting it like that anyway, because they think "yeah yeah technically something something math, I know, but basically in practice it kinda just means that anyway". Even plenty of educated professionals in science do not understand this and continue to get it wrong each day.

Re: Scanning for Pangram Errors

#16
post #12
post #5

Interesting methodology. The results for Pangram are surprisingly good. These tools are definitely not 100% perfect, which is the primary complaint used to dismiss them. However the error rate is also getting impressively low. In most cases I see socially and online, there is a high suspicion that the content is AI generated before someone thinks to submit it to Pangram. It’s used on-demand as a tool to confirm suspi…

I think outside of a few niche uses, e.g. schools/universities, services like Pangram will struggle. Currently people care because so much of AI output is still poor quality. But given I personally know of someone with published AI-generated works where the reviews are gushing and positive across the board, I doubt people will care whether works are AI generated for very long vs whether the text is good . A way of ce…

I actually disagree, I think B people care a lot if a created work is an intended artistic thing vs produced by a machine. It’s like the difference between getting someone a generic hallmark cards vs actually writing something inside. Theres a place for the hallmarks but a plain one with no human touch is still considered tacky

Re: Scanning for Pangram Errors

#17

Pangram is witchcraft to me. The way they can correctly detect AI writing from small samples, with so little statistical signal, is crazy. I've seen it correctly detect AI writing even when people use those "humanizer" rewriting skills that remove the hallmarks of AI writing and make the text indistinguishable from human text. But not to Pangram.

I agree. I ran it on a small sample last night, basically three paragraphs of AI text that I thought read 100% human, except for em dashes, which I use in writing anyway. It flagged the first and last paras as "high" AI and the middle para as "high" human. I have no idea what it is seeing that sets it off.

Re: Scanning for Pangram Errors

#18

Pangram is witchcraft to me. The way they can correctly detect AI writing from small samples, with so little statistical signal, is crazy. I've seen it correctly detect AI writing even when people use those "humanizer" rewriting skills that remove the hallmarks of AI writing and make the text indistinguishable from human text. But not to Pangram.

Consider that they are in effect detecting whether text has been generated by one of a few dozen entities that has produced more text for them to analyze than any given human author.

Now imagine if they put the same effort into detecting if a given text was produced by one of a few dozen super-prolific human writers. I'd imagine they'd get pretty good at that too.

The main limitation of those "humanizer" rewriters is that most of them focus on making the next less detectable to humans, by making them read better. There's likely to be plenty of signal left that isn't affected by trying to make the text read better the same way human writers have plenty of idiosyncrasies despite being human.

Re: Scanning for Pangram Errors

#19
post #8

On the other hand, see https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-i... where Freddie deBoer shows some instances where Pangram is definitely giving wrong or misleading results. (Somewhat-plausible-to-me explanation: It's looking for various stylistic features; older writing very rarely has the most AI-like features, or perhaps almost always has some non-AI-like features that outweigh whatever signs of…

It'd be interesting to see what their detection rate is for AI text made to emulate old books.

Re: Scanning for Pangram Errors

#20
I’m a teacher that requires students to write, so I’m always on the lookout for reliable AI detection. Haven’t anything close to reliable yet, but Pangram definitely seemed to be better than anything else in my (very limited) tests.

That said, it did incorrectly flag something that I personally wrote as AI, which baffled me. I spent a while playing around with the text and feeding it back into Pangram, trying to figure out what the issue was. Turns out, I had a section in the text with three bullet points. I removed the bullets (literally just the bullets themselves, no actual text) and it passed as human. While it seems to perform better than anything else, I’m a bit skeptical of their claimed success rate.

Post reply on HN