Live data from Hacker News

OpenAI shuts down its AI Classifier due to poor accuracy

decrypt.co

61–70 of 292 posts

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#61

Earlier quoted context omitted.

Taking away tools don't seem to me like the best response same way taking away things tends never to be. If the problem is people not using it right, that seems to me like it would be designed wrong for what people need it for. Like if the issue is using it wrong with too little sentences, then put a minimum sentence or something to have that minimum likelihood. Same goes for representing what it means. If people don…

I think it will be increasingly irrelevant what specific process generated a text, for example. Already before genAI people did not in general query into how politicians' speeches were crafted etc.

cool beans. I didn't think about it like that. Could be.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#62
post #36

Earlier quoted context omitted.

A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process. There is no way to know when it’s wrong so you’ll either use it to help justify what you want, or ignore it. Edit: this tool is as reliable as a magic 8-ball

This is an unreasonable standard. Outside of trivial situations, there are no infallible tools.

You're right. After reading what I'd wrote, there should be some reasonable expectations about a tool, such as how accurate it is, or what are the consequences to be wrong.

The AI detection tool fails both as it has a low accuracy and could ruin someones reputation and livelihood. If a tool like this helped you pick out what color socks you're wearing, then it's just as good as asking a magic 8-ball if you should wear the green socks.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#63

Earlier quoted context omitted.

There's also the post going around about how it can (and does) falsely flag human posts as AI output, particularly among some autistic people. About as useful as a polygraph, no?

Both false-positives are as useful as the other one, flagged "human" but actually "LLM" vs flagged "LLM" but actually "human". As long as no one put too much weight on the result, no harm would have been done, in either case. But clearly, people can't stay away from jumping to conclusions based on what a simple-but-incorrect tool says.

Seems a tautology no? “As long as we ignore the results the results don’t matter.”

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#64

I'm glad that they did, although they should obviously done an announcement for it. The amount of people in the ecosystem who thinks it's even possible to detect if something is AI written or not when it's just a couple of sentences is staggering high. And somehow, people in power seems to put their faith in some of these tools that guarantee a certain amount of truthfulness when in reality it's impossible they could…

Even the idea of it is bad, ChatGPT is supposed to write indistinguishably from a human.

The "detector" has extremely little information and the only somewhat reasonable criteria are things like style, where ChatGPT certainly has a particular, but by no means unique writing style. And as it gets better it will (by definition) be better at writing in more varied styles.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#65
post #46

Not to say "I can detect chatgpt" but it sure seems to have a similar way of talking even when I say things like: Talk like a "Millennial male who is obsessed with Zelda, their name is bob zelenski" Now the topic isnt about anything millennial or Zelda related, but I'd think that the language model would select sentence and paragraph phrasing differently. Maybe I need to switch to the API.

I've also noticed that ChatGPT tends to respond to short prompts, especially questions, in a predictable format. There are a few characteristics. First, it tends to print a five-paragraph essay, with an introduction, three main points, and a conclusion. Second, it signposts really well. Each of the body paragraphs is marked with either a bullet point or a number or something else that says "I'm starting a new point."…

I have to admit I'm struggling to tell if this was done ironically, but your comment is exactly a five paragraph essay with an introduction, three main points, and a conclusion.

If so, nice meta-commentary.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#66

Earlier quoted context omitted.

Taking away tools don't seem to me like the best response same way taking away things tends never to be. If the problem is people not using it right, that seems to me like it would be designed wrong for what people need it for. Like if the issue is using it wrong with too little sentences, then put a minimum sentence or something to have that minimum likelihood. Same goes for representing what it means. If people don…

I think it will be increasingly irrelevant what specific process generated a text, for example. Already before genAI people did not in general query into how politicians' speeches were crafted etc.

Indeed or whether math was done in your head, on a calculator or by a computer. Math is math and the agent that represents the result gets the credit and blame.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#67

There is inherent conflict in having both an AI tool business and an AI tool detection business. If the first does a good job, the second fails. And vice versa. (On the other hand, maybe there is a lot of money to be made selling both, to different groups?)

I don't think this follows. If they wanted, they could crypographically bias the sampling to make the output detectable without decreasing capabilities at all.

Only people using it deceptively would be affected. No idea what portion of ChatGPT's users that is, would be very interested to know.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#68
I wonder why we need this very thing of AI generated. It's a luddite view of AI. Much like the need to distinguish between handcrafted versus machined products - is there a real utility to knowing this?

For educators looking at evaluating students, essays and the like - we possibly need different ways of evaluation rather than on written asynchronous content for communicating concepts and ideas.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#69

Earlier quoted context omitted.

Firstly, this tool cannot be made better than it is due to the nature of its construction, it is completely intrinsic. Secondly, as LLM models improve, as they are guaranteed to do, this tool can only become worse as it becomes increasingly difficult to distinguish between human and AI written text.

I don't know about neither of those. How is it intrinsic? What stops detection improving just because AI gets better? Assuming it just doesn't become sentient human replica or something I mean AI like this where it's just a language model thing. Plus that's assuming future stuff you can track in the meanwhile and still don't justify "remove it because people dumb and do bad stuff with tool", that'd only justify remov…

The algorithms are trained on minimizing the difference between what the algorithm produces and what a human produces. The better the algorithms the less the difference. The algorithms are at the point where there is very little difference and it won’t be long until there is no difference.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#70

Earlier quoted context omitted.

There's also the post going around about how it can (and does) falsely flag human posts as AI output, particularly among some autistic people. About as useful as a polygraph, no?

Both false-positives are as useful as the other one, flagged "human" but actually "LLM" vs flagged "LLM" but actually "human". As long as no one put too much weight on the result, no harm would have been done, in either case. But clearly, people can't stay away from jumping to conclusions based on what a simple-but-incorrect tool says.

Flagged "human" but actually "LLM" is not a false positive, but a false negative.
Post reply on HN