Live data from Hacker News

OpenAI shuts down its AI Classifier due to poor accuracy

decrypt.co

101–110 of 292 posts

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#101

I wonder why we need this very thing of AI generated. It's a luddite view of AI. Much like the need to distinguish between handcrafted versus machined products - is there a real utility to knowing this? For educators looking at evaluating students, essays and the like - we possibly need different ways of evaluation rather than on written asynchronous content for communicating concepts and ideas.

>is there a real utility to knowing this?

For civics, I would say yes.

Imagine you were talking to an online group about a design project for a local neighborhood. Based on the plurality of voices it seemed like mist people wanted a brown and orange design. But later when you talk to actual people in real life, you could only find a few that actually wanted that.

Virtual beings are a great addition to the bot nets that generate false consensus.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#103
post #73
post #72

Earlier quoted context omitted.

> The AI tool doesn't give accurate results. Nearly everything doesn't give 100% accurate results. Even CPUs have had bugs their calculation. You have to use a suitable tool for a suitable job with the correct context while understanding it's limitation to apply it correctly. Now that is proper engineering. You're partially correctly but you're overstating: > A tool that gives incorrect and inconsistent results shoul…

The thread is about tools to evaluate LLMs. Please re-read my comment in that light and generously assume I'm talking about that.

Your comment applies to all these tools though lol. No need to clarify, it's all a probabilistic machine that's very unreliable.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#104

I'm in the SEO game and I've spoken with some 'heavy players' who believe a "Google Ai Update" is in the works. As it currently stands, the search engine results will be completely overtaken by Ai content in the near future without this. From my understanding, this is a fools play in the long run, but there are current Ai Classifier Detectors that can successfully detect ChatGPT and other models (Originality.ai being…

Google needs to do a full 180, and only the most succinct website that answers search queries should be elevated. The current state of Google is a disaster, everything is 100 paragraphs per article, the answer you are looking for buried half way in to make sure you spend more time and scroll to appease the algorithm. I cannot wait for them to sink all these spam websites.

Waiting for Google to do that won't happen, they'd lose too many ad links.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#105
post #3

Earlier quoted context omitted.

How would you go about watermarking AI written text?

Make half of the tokens (the AI's "dictionary") slightly more likely. This would not impact output quality much, but it would only work for longish outputs. And the token probability "key" could probsbly be reverse engineered with enough output.

It would be pretty easy to figure out against standard word probability in average datasets. Even then the longer this system runs the more likely it is to pollute its own dataset by people learning to write from gpt itself.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#106

Earlier quoted context omitted.

Can you elaborate on "invisible?" The only invisible character I can imagine is a space. It seems like any other character either isn't invisible or doesn't exist (ie, isnt a character). Additionally, if I copy-paste text like this are the invisible characters preserved? Are there a bunch of extra spaces somewhere?

There's a bunch of different "spaces", one is a "zero-width space" which isn't visible but still gets copied with the text. https://en.wikipedia.org/wiki/Zero-width_space

And the second site students will go to is zerospaceremover.com or whatever will show up to strip the junk.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#107
post #95

Earlier quoted context omitted.

I can see that working for specialized equipment like police body cameras, but if every camera manufacturer in the world needs to manage keys and install them securely into their sensors then there will be leaked keys within weeks.

Makes fake image, hold it in front of camera, click, verified image...

The signed timestamp and location would give that away, but those would have to become not configurable by the user.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#108
post #76

Earlier quoted context omitted.

Nitpick: ChatGPR is supposed to write in a way that is indistinguishable from a human, to another human. That doesn't mean that it can't be distguishable by some other means.

I think for small amounts of text there's no way around it being indistinguishable to a machine and not distinguishable to a human. There just aren't that many combinations of words that still flow well. Furthermore as more and more people use it I think we'll find some humans changing their speech patterns subconsciously more to mimic whatever it does. I imagine with longer text there will be things they'll be able…

I think for this sort of problem it is more productive to think in terms of the amount of text necessary for detection, and how reliable such a detection would be, than a binary can/can't. I think similarly for how "photorealistic" a particular graphics tech is; many techs have already long passed the point where I can tell at 320x200 but they're not necessarily all there yet at 4K.

LLMs clearly pass the single sentence test. If you generate far more text than their window, I'm pretty sure they'd clearly fail as they start getting repetitive or losing track of what they've written. In between, it varies depending on how much text you get to look at. A single paragraph is pretty darned hard. A full essay starts becoming something I'm more confident in my assessment.

It's also worth reminding people that LLMs are more than just "ChatGPT in its standard form". As a human trying to do bot detection sometimes, I've noticed some tells in ChatGPT's "standard voice" which almost everyone is still using, but once people graduate from "Write a blog post about $TOPIC related to $LANGUAGE" to "Write a blog post about $TOPIC related to $LANGUAGE in the style of Ernest Hemmingway" in their prompts it's going to become very difficult to tell by style alone.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#109
post #47
post #4

Earlier quoted context omitted.

If it's generated by a SaaS, the service could sign all output with a public key.

Why is this comment being downvoted? OpenAI can internally keep a "hash" or a "signature" of every output it ever generated. Given a piece of text, they should then be able to trace back to either a specific session (or a set of sessions) through which this text was generated in. Depending on the hit rate and the hashing methods used, they may be able to indicate the likelihood of a piece of text being generated by A…

Why would they want to is my question. A single character change would break it.

Then you have database costs of storing all that data forever.

Moreso, it's only for openAI, I don't think it will be too long before other gpt4 level models are around and won't give two shits about catering to the AI identification police.

Re: OpenAI shuts down its AI Classifier due to poor accuracy

#110

I'm glad that they did, although they should obviously done an announcement for it. The amount of people in the ecosystem who thinks it's even possible to detect if something is AI written or not when it's just a couple of sentences is staggering high. And somehow, people in power seems to put their faith in some of these tools that guarantee a certain amount of truthfulness when in reality it's impossible they could…

Even the idea of it is bad, ChatGPT is supposed to write indistinguishably from a human. The "detector" has extremely little information and the only somewhat reasonable criteria are things like style, where ChatGPT certainly has a particular, but by no means unique writing style. And as it gets better it will (by definition) be better at writing in more varied styles.

I'd challenge this assumption. ChatGPT is supposed to convey information and answer questions in a manner that is intelligible to humans. It doesn't mean it should write indistinguishably from humans. It has a certain manner of prose that (to me) is distinctive and, for lack of a better descriptor, silkier, more anodyne, than most human writing. It should only attempt a distinct style if prompted to.
Post reply on HN