This is an advertisement disguised as a "report".
GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
371–380 of 528 posts
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#372Earlier quoted context omitted.
>Bug free software is possible, ... Mr. Turing and his halting problem would like to politely disagree with this assertion.
You misread the comment and DR Turing's paper. Getting all possible software correct is impossible, clearly. Getting all the software you release is more possible because you can choose not to release the software that it is too hard to prove correct. Not that the suggestion is practical or likely, but your assertion that it is impossible is incorrect.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#373Earlier quoted context omitted.
> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it. It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules! Or these darn kids and their unreliable sources like Wikiped…
One issue with this analogy is that calculators really are precise when used correctly. LLMs are not. I do think they can be used in research but not without careful checking. In my own work I’ve found them most useful as search aids and brainstorming sounding boards.
Of course you are right. It is the same with all tools, calculators included, if you use them improperly you get poor results.
In this case they're stochastic, which isn't something people are used to happening with computers yet. You have to understand that and learn how to use them or you will get poor results.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#374Earlier quoted context omitted.
> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it. It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules! Or these darn kids and their unreliable sources like Wikiped…
I doubt that it's common for anyone to read a research paper and then question whether the researcher's calculator was working reliably. Sure, maybe someday LLMs will be able to report facts in a mostly reliable fashion (like a typical calculator), but we're definitely not even close to that yet, so until we are the skepticism is very much warranted. Especially when the details really do matter, as in scientific rese…
LLM's do not work reliably, that's not their purpose.
If you use them that way it's akin to using a butter knife as a screwdriver. You might get away with it once or twice, but then you slip and stab yourself. Better to go find screwdriver if you need reliable.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#375Earlier quoted context omitted.
> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it. It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules! Or these darn kids and their unreliable sources like Wikiped…
One issue with this analogy is that calculators really are precise when used correctly. LLMs are not. I do think they can be used in research but not without careful checking. In my own work I’ve found them most useful as search aids and brainstorming sounding boards.
I made this a separate comment, because it's wildly off topic, but... they actually aren't. Especially for very large numbers or for high precision. When's the last time you did a firmware update on yours?
It's fairly trivial to find lists of calculator flaws and then identify them in research papers. I recall reading a research paper about it in the 00's.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#376I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…
>this error does make me pause to wonder how much of the rest of the paper used AI assistance And this is what's operative here. The error spotted, the entire class of error spotted, is easily checked/verified by a non-domain expert. These are the errors we can confirm readily, with obvious and unmistakable signature of hallucination. If these are the only errors, we are not troubled. However: we do not know if these…
Also everyone I know has been relying on google scholar for 10+ years. Is that AI-ish? There are definitely errors on there. If you would extrapolate from citation issues to the content in the age of LLMs, were you doing so then as well?
It's the age-old debate about spelling/grammar issues in technical work. In my experience it rarely gets to the point that these errors eg from non-native speakers affect my interpretation. Others claim to infer shoddy content.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#377Earlier quoted context omitted.
One issue with this analogy is that calculators really are precise when used correctly. LLMs are not. I do think they can be used in research but not without careful checking. In my own work I’ve found them most useful as search aids and brainstorming sounding boards.
One issue with this analogy is that paper encyclopedias really are precise when used correctly. Wikipedia is not. I do think it can be used in research but not without careful checking. In my own work I've found it most useful as a search aid and for brainstorming. ^ this same comment 10 years ago
This is really just restating what I already said in this thread, but you're right. That's because wikipedia isn't a primary source and was never, ever meant to be. You are SUPPOSED to go read it then click through to the primary sources and cite those.
Lots of people use it incorrectly and get bad results because they still haven't realized this... all these years later.
Same thing with treating stochastic LLM's like sources of truth and knowledge. Those folks are just doing it wrong.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#378The old create the problem and sell the solution shtick.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#379Earlier quoted context omitted.
>this error does make me pause to wonder how much of the rest of the paper used AI assistance And this is what's operative here. The error spotted, the entire class of error spotted, is easily checked/verified by a non-domain expert. These are the errors we can confirm readily, with obvious and unmistakable signature of hallucination. If these are the only errors, we are not troubled. However: we do not know if these…
> If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care. I am troubled by people using an LLM at all to write academic research p…
Google Translate et al were never good enough at this task to actually allow people to use the results for anything professional. Previous tools were limited to getting a rough gloss of what words in another language mean.
But LLMs can be used in this way, and are being used in this way; and this is increasingly allowing non-English-fluent academics to publish papers in English-language journals (thus engaging with the English-language academic community), where previously those academics they may have felt "stuck" publishing in what few journals exist for their discipline in their own language.
Would you call the use of LLMs for translation "shoddy" or "irresponsible"? To me, it'd be no more and no less "shoddy" or "irresponsible" than it would be to hire a freelance human translator to translate the paper for you. (In fact, the human translator might be a worse idea, as LLMs are more likely to understand how to translate the specific academic jargon of your discipline than a randomly-selected human translator would be.)
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#380Earlier quoted context omitted.
> It's a shoddy, irresponsible way to work. And also plagiarism, when you claim authorship of it. It reminds me of kids these days and their fancy calculators! Those new fangled doohickeys just aren't reliable, and the kids never realize that they won't always have a calculator on them! Everyone should just do it the good old fashioned way with slide rules! Or these darn kids and their unreliable sources like Wikiped…
Annoying dismissal. In an academic paper, you condense a lot of thinking and work, into a writeup. Why would you blow off the writeup part, and impose AI slop upon the reviewers and the research community?
They should still review the final result though. There is no excuse for not doing that.