Live data from Hacker News

Detect ChatGPT Generated Content

gptzero.me

101–109 of 109 posts

Re: Detect ChatGPT Generated Content

#101
post #76

Earlier quoted context omitted.

At the university level, the problem of cheating is more a problem with our idea of a university. In an ideal world, you go to a university to learn. If someone wants to cheat, ultimately they're just cheating themselves (and, if at a private university or even most public universities, throwing a lot of money away). "Cheating" isn't really the university's problem in that view. Ofc, in reality universities aren't (o…

Universities should stop issuing transcripts and diplomas that they refuse to guarantee. They are complicit in a fraud against employers.

That's on employers for putting too much faith in educational institutions.

It's like expecting a guarantee from the church that all their exorcisms render the host demon-free.

A degree is an artifact of a belief system. That there are credentialed idiots among us means the belief system itself is flawed.

Re: Detect ChatGPT Generated Content

#102
post #26

The problem with such things is that you would need 100% accuracy in most cases for it to be useful. For example many schools and universities fear that students use ChatGPT for homework. If such a plagiarism checker has false positive results in just a small percentage of cases, the consequences for honest students would be too severe to actually act on the results of the check. But perfect accuracy can never be rea…

At the university level, the problem of cheating is more a problem with our idea of a university. In an ideal world, you go to a university to learn. If someone wants to cheat, ultimately they're just cheating themselves (and, if at a private university or even most public universities, throwing a lot of money away). "Cheating" isn't really the university's problem in that view. Ofc, in reality universities aren't (o…

Grades are also needed in university to check for prerequisite knowledge for courses. You waste everybody's time and resources if you let people into courses who think or pretend that they have the prerequisite knowledge. People who are cheating aren't only cheating themselves.

Re: Detect ChatGPT Generated Content

#103
post #34

Earlier quoted context omitted.

I ran your comment through this Detector, and received the following message. Your text is likely to be written entirely by AI False positives and all.

Now you've said this I read that comment and it has the ChatGPT tone for sure.

Yes, it was ChatGPT indeed.

Re: Detect ChatGPT Generated Content

#104
post #82

We ought to require OpenAI to run something like this: a “Hey ChatGPT, this you?” endpoint that replies Yes/No/similarity score, and a creation timestamp, when you hit it with some text. As regulations go, this one’s not too burdensome: expensive, but pretty cheap compared to training and running a large language model in the first place.

If OpenAI implements such an endpoint, then a big chunk of potential customers would just not use ChatGPT… I mean, if I’m a student planning to use ChatGPT to enhance my uni essays, but it turns out OpenAI actually can say to my uni if my texts were AI generated, well, I’m not gonna use ChatGPT at all.

Re: Detect ChatGPT Generated Content

#105

Earlier quoted context omitted.

It works both ways, but generation is advantaged in the long run. There has to actually be a statistical difference to detect, and AI outputs without statistical differences from human output are obviously possible, since humans make them all the time.

Sort off, I’m aware that in principle the generator has an advantage and eventually the detector will average out to a coin flip at best. However some advantages can disappear when you put constraints on the output such as quality and correctness. So whilst the end result might be less statistically significant in terms of was it human or AI generated it can overall be also less useful to the end user.

"However some advantages can disappear when you put constraints on the output such as quality and correctness."

Only if you suppose that the ideal output is superhuman. In the case of OpenAI et al, that's arguably the case, but those aren't the players that are going to get into an arms race with detection anyway. They want it to be relatively easy to detect AI generated content, because they're not in the plagiarism business, and anti-plagiarism measures will get the public and media off their backs. And nobody who is interested in targeting plagiarism has nearly the funding to build their own LLM on a level that matters.

So if there's an arms race in the near term, I expect it will be with postprocessors instead. These will be much smaller models (i.e. runs in your browser, or at least on a small backend machine) that take the output of ChatGPT and tweak it to fool detectors. They won't care about maximizing quality or accuracy, but will just care about preserving meaning while erasing statistical signs of AI generation.

I don't know if the business case for that will be there. It's there for selling papers, and almost certainly some people will try their hand at these models just for the challenge and/or to prove a point.

Re: Detect ChatGPT Generated Content

#106

Earlier quoted context omitted.

Sort off, I’m aware that in principle the generator has an advantage and eventually the detector will average out to a coin flip at best. However some advantages can disappear when you put constraints on the output such as quality and correctness. So whilst the end result might be less statistically significant in terms of was it human or AI generated it can overall be also less useful to the end user.

"However some advantages can disappear when you put constraints on the output such as quality and correctness." Only if you suppose that the ideal output is superhuman. In the case of OpenAI et al, that's arguably the case, but those aren't the players that are going to get into an arms race with detection anyway. They want it to be relatively easy to detect AI generated content, because they're not in the plagiarism…

Im not sure if that how it actually would work out.

Most humans can’t write say an essay to save their life.

And those who do write very well tend to have their own signature.

Whilst it’s not 100% accurate we’ve managed to fairly successfully attribute a lot of unknown works to specific authors based on their known works.

So if you create a generator that produces output equals to say top 1% of human authors I’m not entirely sure that you can get one that doesn’t have its own signature.

Because whilst as you said most humans produce output that is statistically indistinguishable from most other humans the output that tends to survive selection bias and become known works is quite distinguishable by definition.

So you don’t even need to get to superhuman capability you just need to get to a high enough output quality that it would limit the statistical search space from billions to millions or even thousands.

Re: Detect ChatGPT Generated Content

#107

Earlier quoted context omitted.

"However some advantages can disappear when you put constraints on the output such as quality and correctness." Only if you suppose that the ideal output is superhuman. In the case of OpenAI et al, that's arguably the case, but those aren't the players that are going to get into an arms race with detection anyway. They want it to be relatively easy to detect AI generated content, because they're not in the plagiarism…

Im not sure if that how it actually would work out. Most humans can’t write say an essay to save their life. And those who do write very well tend to have their own signature. Whilst it’s not 100% accurate we’ve managed to fairly successfully attribute a lot of unknown works to specific authors based on their known works. So if you create a generator that produces output equals to say top 1% of human authors I’m not…

This may be along the lines of what you’re suggesting, but what if you flipped this around: instead of trying to recognize AI, you recognize the student? You model each student’s quirks so you can tell if they wrote their essay, or if someone else did. Now you don’t care about AI specifically; you just care about whether they wrote what they submitted.

The main failure mode I see here is students dramatically improving and throwing the system off. If someone gets a tutor or goes to writing workshops, you don’t want to accuse them of plagiarism just because they got better. But there may be ways you could deal with that, like having the student submit new samples.

Re: Detect ChatGPT Generated Content

#108

Earlier quoted context omitted.

Im not sure if that how it actually would work out. Most humans can’t write say an essay to save their life. And those who do write very well tend to have their own signature. Whilst it’s not 100% accurate we’ve managed to fairly successfully attribute a lot of unknown works to specific authors based on their known works. So if you create a generator that produces output equals to say top 1% of human authors I’m not…

This may be along the lines of what you’re suggesting, but what if you flipped this around: instead of trying to recognize AI, you recognize the student? You model each student’s quirks so you can tell if they wrote their essay, or if someone else did. Now you don’t care about AI specifically; you just care about whether they wrote what they submitted. The main failure mode I see here is students dramatically improvi…

That could work but that is changing the problem and moving the goal posts, a plagiarism detection system that is essentially trained on individual authors would be able to identify any time they skew too far from their rolling average.

I’m not even sure if ML is absolutely necessary for this or not.

Post reply on HN