Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

181–190 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#181

Which is worse: a) p-hacking and suppressing null results b) hallucinations c) falsifying data Would be cool to see an analysis of this

I'm doing some research, and this is something I'm unsure of. I see that "suppressing null results" is a bad thing, and I sort of agree, but for me personally, a lot of the null results are just the result of my own incompetence and don't contain any novel insights.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#182

Getting papers published is now more about embellishing your CV versus a sincere desire to present new research. I see this everywhere at every level. Getting a paper published anywhere is a checkbox in completing your resume. As an industry we need to stop taking this into consideration when reviewing candidates or deciding pay. In some sense it has become an anti-signal.

I think its fairer to say that perverse incentives have added more noise to the publishing signal. Publishing 0 times is not better than 100 times, even if 90% of those are Nth author formality/politeness citations.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#183
post #26

This suggests that nobody was screening this papers in the first place—so is it actually significant that people are using LLMs in a setting without meaningful oversight? These clearly aren't being peer-reviewed, so there's no natural check on LLM usage (which is different than what we see in work published in journals).

As one who reviews 20+ papers per year, we don't have time to verify each reference. We verify: is the stuff correct, and is it worthy of publication (in the given venue) given that it is correct. There is still some trust in the authors to not submit made-up-stuff, albeit it is diminishing.

Sorry, but if someone makes a claim and cites a reference, how do you verify "is the stuff correct" without checking that reference?

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#184

Earlier quoted context omitted.

> I'd love to see future reporting that instead of saying "Research finds amazing chemical x which does y" you see "Researcher reproduces amazing results for chemical x which does y. First discovered by z". Most people (that I talk to, at least) in science agree that there's a reproducibility crisis. The challenge is there really isn't a good way to incentivize that work. Fundamentally (unless you're independent weal…

> Eventually you may arrive at something like the H-index, which is defined as "The highest number H you can pick, where H is the number of papers you have written with H citations." It's the Google search algorithm all over again. And it's the certificate trust hierarchy all over again. We keep working on the same problems. Like the two cases I mentioned, this is a matter of making adjustments until you have the des…

Incentives.

First X people that reproduce Y get Z percent of patent revenue.

Or something similar.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#185

Earlier quoted context omitted.

In my mental model, the fundamental problem of reproducibility is that scientists have very hard time to find a penny to fund such research. No one wants to grant “hey I need $1m and 2 years to validate the paper from last year which looks suspicious”. Until we can change how we fund science on the fundamental level; how we assign grants — it will be indeed very hard problem to deal with.

In theory, asking grad students and early career folks to run replications would be a great training tool. But the problem isn’t just funding, it’s time. Successfully running a replication doesn’t get you a publication to help your career.

Yeah, but doesn't publishing an easily falsifiable paper end one?

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#186

Earlier quoted context omitted.

> the content of the papers themselves are not necessarily invalidated. For example, authors may have given an LLM a partial description of a citation and asked the LLM to produce bibtex (a formatted reference) Maybe I'm overreacting, but this feels like an insanely biased response. They found the one potentially innocuous reason and latched onto that as a way to hand-wave the entire problem away. Science already had…

Isn't disqualifying X months of potentially great research due to a misformed, but existing reference harsh? I don't think they'd be okay with references that are actually made up.

Science relies on trust.. a lot. So things which show dishonesty are penalised greatly. If we were to remove trust then peer reviewing a paper might take months of work or even years.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#187

Earlier quoted context omitted.

> The challenge is there really isn't a good way to incentivize that work. Ban publication of any research that hasn't been reproduced.

> Ban publication of any research that hasn't been reproduced. Unless it is published, nobody will know about it and thus nobody will try to reproduce it.

Just have a new journal of only papers that have been reproduced, and include the reproduction papers.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#188
post #140

Earlier quoted context omitted.

Why not run every submitted paper through GPTZero (before sending to reviewers) and summarily reject any paper with a hallucination?

That's how GPTZero wants to situate themselves. Who would pay them? Conference organizers are already unpaid and undestaffed, and most conferences aren't profitable. I think rejections shouldn't be automatic. Sometimes there are just typos. Sometimes authors don't understand BibTeX. This needs to be done in a way that reduces the workload for reviewers. One way of doing this would be for GPTZero to annotate each pape…

Most publication venues already pay for a plagiarism detection service, it seems it would be trivial to add it on as a cost. Especially given APCs for journals are several thousand dollars, what's a few dollars more per paper.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#189
post #2

Yuck, this is going to really harm scientific research. There is already a problem with papers falsifying data/samples/etc, LLMs being able to put out plausible papers is just going to make it worse. On the bright side, maybe this will get the scientific community and science journalists to finally take reproducibility more seriously. I'd love to see future reporting that instead of saying "Research finds amazing che…

I'd need to see the same scrutiny applied to pre-AI papers. If a field has a poor replication rate, meaning there's a good chance that a given published paper is just so much junk science, is that better or worse than letting AI hallucinate the data in the first place?

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#190
post #26

Earlier quoted context omitted.

As one who reviews 20+ papers per year, we don't have time to verify each reference. We verify: is the stuff correct, and is it worthy of publication (in the given venue) given that it is correct. There is still some trust in the authors to not submit made-up-stuff, albeit it is diminishing.

Sorry, but if someone makes a claim and cites a reference, how do you verify "is the stuff correct" without checking that reference?

Those are typically things you are familiar with or can easily check.

Fake references are more common in the introduction where you list relevant material to strengthen your results. They often don't change the validity of the claim, but the potential impact or value.

Post reply on HN