Live data from Hacker News

OpenAI shuts down its AI Classifier due to poor accuracy

decrypt.co

71–80 of 292 posts

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#71

I'm in the SEO game and I've spoken with some 'heavy players' who believe a "Google Ai Update" is in the works. As it currently stands, the search engine results will be completely overtaken by Ai content in the near future without this. From my understanding, this is a fools play in the long run, but there are current Ai Classifier Detectors that can successfully detect ChatGPT and other models (Originality.ai being…

Google needs to do a full 180, and only the most succinct website that answers search queries should be elevated.

The current state of Google is a disaster, everything is 100 paragraphs per article, the answer you are looking for buried half way in to make sure you spend more time and scroll to appease the algorithm.

I cannot wait for them to sink all these spam websites.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#72
post #57
post #49

Earlier quoted context omitted.

> A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process. It can be used for some decision (i.e. not critical ones), but it should NOT be used to accused someone of academic misconduct unless the tool meets a very robust quality standard. > this tool is as reliable as a magic 8-ball Citation needed

The AI tool doesn't give accurate results. You don't know when it's not accurate. There is no accurate way to check its results. Who should use a tool to help them make a decision when you don't know when the tool will be wrong and it has a low rate of accuracy? It's in the article.

> The AI tool doesn't give accurate results.

Nearly everything doesn't give 100% accurate results. Even CPUs have had bugs their calculation. You have to use a suitable tool for a suitable job with the correct context while understanding it's limitation to apply it correctly. Now that is proper engineering. You're partially correctly but you're overstating:

> A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process.

That's totally wrong and an overstated position.

A better position is that some tools have such a low accuracy rate that they shouldn't be used for their intended purpose. Now that position I agree with it. I accept that CPUs may give incorrect results due to a cosmic ray event, but I wouldn't accept a CPU that gives the wrong result for 1/100 instructions.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#73
post #72
post #57

Earlier quoted context omitted.

The AI tool doesn't give accurate results. You don't know when it's not accurate. There is no accurate way to check its results. Who should use a tool to help them make a decision when you don't know when the tool will be wrong and it has a low rate of accuracy? It's in the article.

> The AI tool doesn't give accurate results. Nearly everything doesn't give 100% accurate results. Even CPUs have had bugs their calculation. You have to use a suitable tool for a suitable job with the correct context while understanding it's limitation to apply it correctly. Now that is proper engineering. You're partially correctly but you're overstating: > A tool that gives incorrect and inconsistent results shoul…

The thread is about tools to evaluate LLMs. Please re-read my comment in that light and generously assume I'm talking about that.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#74
post #51

"Half a year later, that tool is dead, killed because it couldn’t do what it was designed to do." This was my conclusion as well testing the image detectors. Current automated detection isn’t very reliable. I tried out Optic’s AI or Not , which boasts 95% accuracy, on a small sample of my own images. It correctly labeled those with AI content as AI generated, but it also labeled about 50% of my own stock photo compos…

> but it also labeled about 50% of my own stock photo composites I tried as AI generated

Could it be that a large proportion of the source stock photos were actually AI generated?

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#75

I'm glad that they did, although they should obviously done an announcement for it. The amount of people in the ecosystem who thinks it's even possible to detect if something is AI written or not when it's just a couple of sentences is staggering high. And somehow, people in power seems to put their faith in some of these tools that guarantee a certain amount of truthfulness when in reality it's impossible they could…

I'm still interested in this line of enquiry.

These models are clearly not good enough for decision-making, but still might tell an interesting story.

Here's an easily testable exercise: get a load of news from somewhere like newsapi.ai, run it through an open model and there should be a clear discontinuity around ChatGPT launch.

We can assume false positives and false negatives, but with a fat wadge of data we should still be able to discern trends.

Certainly couldn't accuse a student of cheating with it, but maybe spot content farms.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#76

I'm glad that they did, although they should obviously done an announcement for it. The amount of people in the ecosystem who thinks it's even possible to detect if something is AI written or not when it's just a couple of sentences is staggering high. And somehow, people in power seems to put their faith in some of these tools that guarantee a certain amount of truthfulness when in reality it's impossible they could…

Even the idea of it is bad, ChatGPT is supposed to write indistinguishably from a human. The "detector" has extremely little information and the only somewhat reasonable criteria are things like style, where ChatGPT certainly has a particular, but by no means unique writing style. And as it gets better it will (by definition) be better at writing in more varied styles.

Nitpick: ChatGPR is supposed to write in a way that is indistinguishable from a human, to another human.

That doesn't mean that it can't be distguishable by some other means.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#77
post #49
post #36

Earlier quoted context omitted.

A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process. There is no way to know when it’s wrong so you’ll either use it to help justify what you want, or ignore it. Edit: this tool is as reliable as a magic 8-ball

> A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process. It can be used for some decision (i.e. not critical ones), but it should NOT be used to accused someone of academic misconduct unless the tool meets a very robust quality standard. > this tool is as reliable as a magic 8-ball Citation needed

>"should NOT be used to accused someone of academic misconduct unless the tool meets a very robust quality standard."

Meanwhile, the leading commercial tools for plagiarism detection often flag properly cited/annotated quotes from sources in your text as plagiarism.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#78

Earlier quoted context omitted.

All that has to be shown is that the tool is as bad as or worse than random today , in order to remove it today.

From the article, "while incorrectly labeling the human-written text as AI-written 9% of the time." Seems like from what the article we're talkin about says it definitely ain't worse than random by far. Thing you most want to avoid is wrongly labeling humans as AI-written so that seems pretty good. Though it only identified 26% of AI text as "likely AI-written" that's still better than nothing, and better than random…

You're right, I should have been less specific. If the harm of false positives is significant you may not need to have random or worse than random results to feel obligated to stop the project.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#79
post #74
post #51

"Half a year later, that tool is dead, killed because it couldn’t do what it was designed to do." This was my conclusion as well testing the image detectors. Current automated detection isn’t very reliable. I tried out Optic’s AI or Not , which boasts 95% accuracy, on a small sample of my own images. It correctly labeled those with AI content as AI generated, but it also labeled about 50% of my own stock photo compos…

> but it also labeled about 50% of my own stock photo composites I tried as AI generated Could it be that a large proportion of the source stock photos were actually AI generated?

No, they were older images. However, that is now becoming a problem. Some stock photo sites now have AI images and they are not labeled. I'm able to distinguish most for now because at hires the details contain obvious errors.

This is really painful, because for some of my work I need high quality images suitable for print. Now I can't just look at the thumbnail and say "this will work". I now have to examine it taking more of my time.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#80
post #76

Earlier quoted context omitted.

Even the idea of it is bad, ChatGPT is supposed to write indistinguishably from a human. The "detector" has extremely little information and the only somewhat reasonable criteria are things like style, where ChatGPT certainly has a particular, but by no means unique writing style. And as it gets better it will (by definition) be better at writing in more varied styles.

Nitpick: ChatGPR is supposed to write in a way that is indistinguishable from a human, to another human. That doesn't mean that it can't be distguishable by some other means.

I think for small amounts of text there's no way around it being indistinguishable to a machine and not distinguishable to a human. There just aren't that many combinations of words that still flow well. Furthermore as more and more people use it I think we'll find some humans changing their speech patterns subconsciously more to mimic whatever it does. I imagine with longer text there will be things they'll be able to find, but, I think it will end up being trivial for others to detect what those changes are and then modifying the result enough to be undetectable.
Post reply on HN