Live data from Hacker News

OpenAI shuts down its AI Classifier due to poor accuracy

decrypt.co

41–50 of 292 posts

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#41

Earlier quoted context omitted.

There are numerous people that I’ve tried to get them comprehend statistics, important medical statistics for doctors so you would assume they’re smart enough to understand. There just seems to be a sufficient subset of the population that are blind to statistics and nothing can be done about it. Even sitting down and carefully going through the math with them doesn’t work. No matter how deep into visualization rabbi…

Alright let's say that's how it is. How happy would everyone else be if they were treated like that even if they weren't like that? I'd be right miffed and I ain't no einstein. My problem is saying it's a good thing to *remove* options just because some people don't know how to use it. Use that kinda logic for other stuff and you'd paint yourself in a corner with a very angry hornet trapped in it, so not the kind of…

What about the patients getting unnecessary treatments? How upset should they be? What about the student expelled for AI plagiarism due to a false reading? These things are unreliable, and despite an infinite amount of caveats there is no way to prevent people from over relying on it. We might as well dunk people in the water to see if they float.

That’s a weird kind of extortion, a demand that we placate a subset of the population to the detriment of others. If a conflict came down to people who understand stats versus those blind to it I would put my money on those who understand stats.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#42

Earlier quoted context omitted.

I don't agree with that premise. I don't know that it can't work, that'd suggest something like no matter what it's worse than a coin flip. I don't think it's that bad or at least nobody showed me anything of it being that bad. You'd have to show me that it can't work and that seems to me a pretty big ask I know

All that has to be shown is that the tool is as bad as or worse than random today , in order to remove it today.

From the article, "while incorrectly labeling the human-written text as AI-written 9% of the time."

Seems like from what the article we're talkin about says it definitely ain't worse than random by far. Thing you most want to avoid is wrongly labeling humans as AI-written so that seems pretty good. Though it only identified 26% of AI text as "likely AI-written" that's still better than nothing, and better than random. But we don't know or I don't know from the article if that's on the problem cases of less than 1,000 characters or not. It don't say what the *best case* is just what the general cases are.

Anyhow don't seem to me worse than random is the issue here

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#43
post #36

Earlier quoted context omitted.

Both false-positives are as useful as the other one, flagged "human" but actually "LLM" vs flagged "LLM" but actually "human". As long as no one put too much weight on the result, no harm would have been done, in either case. But clearly, people can't stay away from jumping to conclusions based on what a simple-but-incorrect tool says.

A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process. There is no way to know when it’s wrong so you’ll either use it to help justify what you want, or ignore it. Edit: this tool is as reliable as a magic 8-ball

If you were trying to predict the direction a stock will move (up or down) and it was right 99.9% of the time, would you use it or not?

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#44

Earlier quoted context omitted.

Alright let's say that's how it is. How happy would everyone else be if they were treated like that even if they weren't like that? I'd be right miffed and I ain't no einstein. My problem is saying it's a good thing to *remove* options just because some people don't know how to use it. Use that kinda logic for other stuff and you'd paint yourself in a corner with a very angry hornet trapped in it, so not the kind of…

What about the patients getting unnecessary treatments? How upset should they be? What about the student expelled for AI plagiarism due to a false reading? These things are unreliable, and despite an infinite amount of caveats there is no way to prevent people from over relying on it. We might as well dunk people in the water to see if they float. That’s a weird kind of extortion, a demand that we placate a subset of…

I don't see how that's any different from anything, any tool, any power, any method. Same problem with everything. That's why this don't convince me and just seems like removing things cynically instead of improving it. Seems to me like the company also really don't want its service identified negatively like that and get itself associated with cheaters even if they're the ones selling the cheat identifying, or something like that.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#45

Earlier quoted context omitted.

We could combine those, couldn't we?

You could but is there any reason to believe these two noisy signals wouldn't result in more combined noise than signal? Sure, it's theoretically possible to add two noisy signals that are uncorrelated and get noise reduction, but is it probable this would be such a case?

Yes, you can :)

It all depends on the properties of the signal and the noise. In photography you can combine multiple noisy images to increase the signal to noise ratio. This works because the signal increases O(N) with the number of images but the noise only increases O(sqrt(N)). The result is that while both signal and noise are increasing, the signal is increasing faster.

I have no idea if this idea could be used for AI detection, but it is possible to combine 2 noisy signals and get better SNR.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#46

Not to say "I can detect chatgpt" but it sure seems to have a similar way of talking even when I say things like: Talk like a "Millennial male who is obsessed with Zelda, their name is bob zelenski" Now the topic isnt about anything millennial or Zelda related, but I'd think that the language model would select sentence and paragraph phrasing differently. Maybe I need to switch to the API.

I've also noticed that ChatGPT tends to respond to short prompts, especially questions, in a predictable format. There are a few characteristics.

First, it tends to print a five-paragraph essay, with an introduction, three main points, and a conclusion.

Second, it signposts really well. Each of the body paragraphs is marked with either a bullet point or a number or something else that says "I'm starting a new point."

Third, it always reads like a WikiHow article. There's never any subtle humour or self-deprecation or ironic understatement. It's very straightforward, like an infographic.

It's definitely easy to recognize a ChatGPT response to a simple prompt if the author hasn't taken any measures to disguise it. The conclusion usually has a generic reminder that your mileage may vary and that you should always be careful.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#47
post #4
post #3

Earlier quoted context omitted.

How would you go about watermarking AI written text?

If it's generated by a SaaS, the service could sign all output with a public key.

Why is this comment being downvoted?

OpenAI can internally keep a "hash" or a "signature" of every output it ever generated.

Given a piece of text, they should then be able to trace back to either a specific session (or a set of sessions) through which this text was generated in.

Depending on the hit rate and the hashing methods used, they may be able to indicate the likelihood of a piece of text being generated by AI.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#48
I'm in the SEO game and I've spoken with some 'heavy players' who believe a "Google Ai Update" is in the works. As it currently stands, the search engine results will be completely overtaken by Ai content in the near future without this.

From my understanding, this is a fools play in the long run, but there are current Ai Classifier Detectors that can successfully detect ChatGPT and other models (Originality.ai being a big one) on longish content.

Their process is fairly simple, they create a classification model after generating tons of examples from all the major models (ChatGPT, GPT4, Laama, etc).

One obvious downside to their strategy is the implementation of Finetuning and how that changes the stylistic output. This same 'heavy hitter' has successfully bypassed Originalities detector using his specified finetuning method (which he said took months of testing and thousands of dollars).

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#49
post #36

Earlier quoted context omitted.

Both false-positives are as useful as the other one, flagged "human" but actually "LLM" vs flagged "LLM" but actually "human". As long as no one put too much weight on the result, no harm would have been done, in either case. But clearly, people can't stay away from jumping to conclusions based on what a simple-but-incorrect tool says.

A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process. There is no way to know when it’s wrong so you’ll either use it to help justify what you want, or ignore it. Edit: this tool is as reliable as a magic 8-ball

> A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process.

It can be used for some decision (i.e. not critical ones), but it should NOT be used to accused someone of academic misconduct unless the tool meets a very robust quality standard.

> this tool is as reliable as a magic 8-ball

Citation needed

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#50

Earlier quoted context omitted.

We could combine those, couldn't we?

You could but is there any reason to believe these two noisy signals wouldn't result in more combined noise than signal? Sure, it's theoretically possible to add two noisy signals that are uncorrelated and get noise reduction, but is it probable this would be such a case?

If the noisy signals are not completely correlated then the signal would be enhanced; however in this case I imagine that there is likely to be a strong correlation between different tools which would mean adding additional sources may not be so useful.
Post reply on HN