Live data from Hacker News

GPTZero Case Study – Exploring False Positives

gonzoknows.com

1–10 of 120 posts

Re: GPTZero Case Study – Exploring False Positives

#6
post #2

Alternative title: Most academic papers are indistinguishable from AI generated babble.

...to AI. It's kinda funny how this is yet another area where these models suck very much in the same way that most humans do. LLMs are bad at arithmetic? So are most people. Can't tell science from babble? I already wouldn't ask a non-expert to rate any aspect of an academic paper. Trusting the average Joe who has only completed some basic form of education would be tremendously stupid. Same with these models. Maybe we can get more out of it in specific areas with fine tuning, but we're very far away from a universal expert system.

Re: GPTZero Case Study – Exploring False Positives

#8
I’ve got a fun little side project that uses GPT. I tested gptzero against 10 of my projects’ writings and 10 of my own. It detected 6 out of 10 correctly in both cases (4 gpt-written bits were declared human, 4 human-written were declared gpt).

Which is better than 50% but not nearly good enough to base any kind of decision on.

Re: GPTZero Case Study – Exploring False Positives

#10
post #2

Alternative title: Most academic papers are indistinguishable from AI generated babble.

...to AI. It's kinda funny how this is yet another area where these models suck very much in the same way that most humans do. LLMs are bad at arithmetic? So are most people. Can't tell science from babble? I already wouldn't ask a non-expert to rate any aspect of an academic paper. Trusting the average Joe who has only completed some basic form of education would be tremendously stupid. Same with these models. Maybe…

The best was the Ted Chiang article making numerous category errors and forest/trees mistakes in arguing that LLMs just store lossy copies of their training data. It was well-written, plausible, and so very incorrect.
Post reply on HN