Live data from Hacker News

OpenAI shuts down its AI Classifier due to poor accuracy

decrypt.co

221–230 of 292 posts

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#221

I'm glad that they did, although they should obviously done an announcement for it. The amount of people in the ecosystem who thinks it's even possible to detect if something is AI written or not when it's just a couple of sentences is staggering high. And somehow, people in power seems to put their faith in some of these tools that guarantee a certain amount of truthfulness when in reality it's impossible they could…

Based on my experience from grad school, I would bet plenty of the professors who fail students because ChatGPT said ChatGPT might have written something honestly don't care whether it's true or not, as long as it shifts liability away from themselves onto someone else

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#222
post #55

Earlier quoted context omitted.

If you were trying to predict the direction a stock will move (up or down) and it was right 99.9% of the time, would you use it or not?

This is a strawman. First, the AI detection algorithms can't offer anything close to 99.9%. Second, your scenario doesn't analyze another human and issue judgement, as the AI detection algorithms do. When a human is miscategorized as a bot, they could find themselves in front of academic fraud boards, skipped over by recruiters, placed in the spam folder, etc.

> Second, your scenario doesn't analyze another human and issue judgement, as the AI detection algorithms do.

> When a human is miscategorized as a bot, they could find themselves in front of academic fraud boards, skipped over by recruiters, placed in the spam folder, etc.

Is the problem here the algorithms or how people choose to use them?

There’s a big difference between treating the results of an AI algorithm as infallible, and treating it as just one piece of probabilistic evidence, to be combined with others, to produce a probabilistic conclusion.

“AI detector says AI wrote student’s essay, therefore it must be true, so let’s fail/expel/etc them” vs “AI detector says AI wrote student’s essay, plus I have other independent reasons to suspect that, so I’m going to investigate the matter further”

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#223
post #195

Earlier quoted context omitted.

There are already quite complex language models which can run on a CPU. Outside of the government banning personal LLMs, the chance of there not existing a working fully FOSS and open data rewrite model, if it becomes known that ChatGPT output is marked, seems very low. The water marking techniques also can not work after some level of sophisticated rewriting. There simply will be no data encoded in the probabilities…

If it's sophisticatedly rewritten then it's no longer AI generated

That is not a reliable indicator even today. GPT-4 (not the ChatGPT RLHF one) is not distinguishable from human writing. You could ask it about modern events, but that's not a long term plan, and it could just make the excuse they don't follow the news.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#224

Earlier quoted context omitted.

Polygraph is pseudo-science, it measures nothing.

I mean, it literally and factually measures multiple your body's autonomous responses - all of which are provably correlated with stress. That's what a polygraph machine is . Saying it measures nothing is factually incorrect. You can't detect "truth" from that, but you can often tell (i.e. with better accuracy than chance) whether or not a subject is able to give a confident, uncomplicated yes-or-no to a straightforw…

Right. The point is: it absolutely does NOT measure what it claims to measure, i.e. truthfulness.

You can detect indicators of stress... or hot weather... or stage-fright (admittedly a form of stress)... or too much caffeine... or an underlying (maybe undiagnosed) medical condition, etc. So it does not even necessarily measure "stress".

It's about as useful as the so called "fruit machine" which they used to test for homosexuality[0], in that it is utterly useless while at the same time can be quite ruinous for people. People have been fired over polygraph "fails", and while not admissible in courts, people probably have been fingered for crimes after they failed polygraphs. Also, criminals have gone free after passing polygraphs[1].

>But everyone knows that it's not very reliable in almost every circumstance it's used.

You and I may know that. But a lot of people actually do not. That's why it's still used. Either because people administering those tests think it's "good science", or because those people administering it know that while it's all bullshit the person they are testing might not know that and break down and admit to things. Remember that fake polygraph on the show The Wire, which was just a copier they strapped to the suspect. If I remember correctly that was based upon true events.

A quick google shows e.g. you can hire "polygraphers" to e.g. "test" if your partner was unfaithful, making claims such as: "However, assuming that you have a good polygrapher with a fair amount of experience in working with betrayal trauma, you're going to get results that are at least 90% accurate or better."[2]

The US (and probably a lot of other) government(s) like their polygraphs very much, too[3].

> you can often tell (i.e. with better accuracy than chance) whether or not a subject is able to give a confident, uncomplicated yes-or-no to a straightforward question in a situation where they don't have to be particularly nervous

Uhmm, if somebody sat me down in a room, strapped all kinds of "science" to my body and then asked me questions, I'd be quite nervous regardless of whether I am truthful or not. In fact, I'd be even more nervous knowing it's a polygraph and bullshit, because I cannot know if the person administrating it would know that too.

If that somebody then asked me "Have you ever killed a prostitute?", or "Have you ever colluded with the enemy?", or "Have you ever cheated on your partner?", or "Have you ever stolen from your employer?", for example, my stress would certainly peak despite being able to confidently and truthfully answer "No!" to all of those questions. And I am sure the polygraph would "measure" my "stress".

[0] Yes, that was a real thing too. https://en.wikipedia.org/wiki/Fruit_machine_(homosexuality_t...

[1] E.g. the Green River Killer Gary Ridgway passed a polygraph, so the police turned their resources to another suspect who failed the polygraph. That was in 1984. Ridgway remained free until his arrest in 2001. He killed at least 4 more times after the investigation stopped focusing on him after that "passed" polygraph.

[2] https://www.affairrecovery.com/newsletter/founder/use-abuse-...

[3] https://support.clearancejobs.com/t/the-differences-between-...

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#225
I urge anyone who values data privacy to refuse to use sites such as this which employ third party data services, ie. popups, which purport to allow you to "Manage" your consent but which hide a long list of opt-out "Legitimate Interest" flags in a "Vendors" list concealed at the bottom of another scrolling list.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#226
I think educators got away for too long of handling the problem of umotivated students and uninteresting classes subjects...

Not the educators fault tough, more like the system is bad.

My point is that given knowledge is mostly free and available, the system should teach the students to think rather than using tools or remembering facts

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#227

Earlier quoted context omitted.

> I don’t see LLM’s breaking this approach to classify human writing. Why not? Record a bunch of humans writing, train model, release. That's orders of magnitude simpler than to come up with the right text to begin with.

Lol. I love HN -- the reaction is because this is either straight-faced or tongue-in-cheek, if it's straight-faced, this is stylistically a parody of the infamous "well Dropbox is rsync, it's moat is basically a SWE-weekend" comment

Someone can release this as a product though, exactly like Dropbox. Dropbox basically was rsync, it just had a better UX. There are a lot of people these days that are pretty good at taking ML models and slapping a nice UX on top of them.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#228

Earlier quoted context omitted.

Retyping the essay from the chatgpt while actively rewording the occassional sentence seems like it would do it.

It seems like that's nearing the sweet spot of fraud prevention, where committing the act of fraud is as much work as doing the real thing.

Not when someone creates a browser extension that you run while using Google Docs, feed the entire ChatGPT doc into it and it recreates the doc over a 2 hour period with small bits and pieces.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#229
post #58

Humorously, in my experience, if a response from ChatGPT ever got classified as AI generated by tools like ZeroGPT or similar, all I had to do was adjust the prompt to tell the model not to sound like it was AI generated and that bypassed all detection with a very high success rate. Additionally, I also found that if you prompt it to make the response be in the style or some known writer for example, it often made re…

This didn't work half a year ago. They'd still accuse the rewritten text to be AI generated. I think some of the recent updates have changed the tone of ChatGPT so significantly that they no longer register on the radar.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#230
post #187

I'm glad that they did, although they should obviously done an announcement for it. The amount of people in the ecosystem who thinks it's even possible to detect if something is AI written or not when it's just a couple of sentences is staggering high. And somehow, people in power seems to put their faith in some of these tools that guarantee a certain amount of truthfulness when in reality it's impossible they could…

Indeed it's not possible. Say you had a classifier that detected whether a given text was AI generated or not. You can easily plug this classifier into the end of a generative network trying to fool it, and even backpropagate all the way from the yes/no output to the input layer of the generative network. Now you can easily generate text that fools that classifier. So such a model is doomed from the start, unless its…

You just outlined an excellent proof of what might be called the AI Halting Problem.
Post reply on HN