Live data from Hacker News

GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

gptzero.me

201–210 of 528 posts

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#201

Earlier quoted context omitted.

Isn't disqualifying X months of potentially great research due to a misformed, but existing reference harsh? I don't think they'd be okay with references that are actually made up.

When your entire job is confirming that science is valid, I expect a little more humility when it turns out you've missed a critical aspect. How did these 100 sources even get through the validation process? > Isn't disqualifying X months of potentially great research due to a misformed, but existing reference harsh? It will serve as a reminder not to cut any corners.

> When your entire job is confirming that science is valid, I expect a little more humility when it turns out you've missed a critical aspect.

I wouldn't call a misformed reference a critical issue, it happens. That's why we have peer reviews. I would contend drawing superficially valid conclusions from studies through use of AI is a much more burning problem that speaks more to the integrity of the author.

> It will serve as a reminder not to cut any corners.

Or yet another reason to ditch academic work for industry. I doubt the rise of scientific AI tools like AlphaXiv [1], whether you consider them beneficial or detrimental, can be avoided - calling for a level pragmatism.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#202

Earlier quoted context omitted.

In my mental model, the fundamental problem of reproducibility is that scientists have very hard time to find a penny to fund such research. No one wants to grant “hey I need $1m and 2 years to validate the paper from last year which looks suspicious”. Until we can change how we fund science on the fundamental level; how we assign grants — it will be indeed very hard problem to deal with.

In theory, asking grad students and early career folks to run replications would be a great training tool. But the problem isn’t just funding, it’s time. Successfully running a replication doesn’t get you a publication to help your career.

Grad students don’t get to publish a thesis on reproduction. Everyone from the undergraduate research assistant to the tenured professor with research chairs are hyper focused on “publishing” as much “positive result” on “novel” work as possible

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#203
post #2

Yuck, this is going to really harm scientific research. There is already a problem with papers falsifying data/samples/etc, LLMs being able to put out plausible papers is just going to make it worse. On the bright side, maybe this will get the scientific community and science journalists to finally take reproducibility more seriously. I'd love to see future reporting that instead of saying "Research finds amazing che…

In my mental model, the fundamental problem of reproducibility is that scientists have very hard time to find a penny to fund such research. No one wants to grant “hey I need $1m and 2 years to validate the paper from last year which looks suspicious”. Until we can change how we fund science on the fundamental level; how we assign grants — it will be indeed very hard problem to deal with.

Partially. There's also the issue that some sciences, like biology, are a lot messier & less predicatble than people like to believe.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#204

Earlier quoted context omitted.

> Eventually you may arrive at something like the H-index, which is defined as "The highest number H you can pick, where H is the number of papers you have written with H citations." It's the Google search algorithm all over again. And it's the certificate trust hierarchy all over again. We keep working on the same problems. Like the two cases I mentioned, this is a matter of making adjustments until you have the des…

Incentives. First X people that reproduce Y get Z percent of patent revenue. Or something similar.

Patent revenue is mostly irrelevant, as it's too unpredictable and typically decades in the future. Academics rarely do research that can be expected to produce economic value in the next 10–20 years, because the industry can easily outspend the academia in such topics.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#205
post #2

Yuck, this is going to really harm scientific research. There is already a problem with papers falsifying data/samples/etc, LLMs being able to put out plausible papers is just going to make it worse. On the bright side, maybe this will get the scientific community and science journalists to finally take reproducibility more seriously. I'd love to see future reporting that instead of saying "Research finds amazing che…

On the bright side, an LLM can really help set up a reproduction environment.

Perhaps repro should become the basis of peer review?

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#206
post #152
post #133

Earlier quoted context omitted.

usually you reproduce previous research as a byproduct of doing something novel "on top" of the previous result. I dont really see the problem with the current setup. sometimes you can just do something new and assume the previous result, but thats more the exception. youre almost always going to at least in part reproducr the previous one. and if issues come up, its often evident. thats why citations work as a good…

It's often quite common to see a citation say "BTW, we weren't able to reproduce X's numbers, but we got fairly close number Y, so Table 1 includes that one next to an asterisk." The difficult part is surfacing that information to readers of the original paper. The semantic scholar people are beginning to do some work in this area.

yeah thats a good point. the citation might actually be pointing out a problem and not be a point in favor. its a slog to figure out... but seems like the exact type of problem an LLM could handle

give it a published paper and it runs through papers that have cited it and give you an evaluation

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#207

Earlier quoted context omitted.

Reproducibility is overrated and if you could wave a wand to make all papers reproducible tomorrow, it wouldn't fix the problem. It might even make it worse. https://blog.plan99.net/replication-studies-cant-fix-science...

? More samples reduces the variance of a statistic. Obviously it cannot identify systematic bias in a model, or establish causality, or make a "bad" question "good". Its not overrated though -- it would strengthen or weaken the case for many papers.

If you have a strong grip on exactly what it means, sure, but look at any HN thread on the topic of fraud in science. People think replication = validity because it's been described as the replication crisis for the last 15 years. And that's the best case!

Funding replication studies in the current environment would just lead to lots of invalid papers being promoted as "fully replicated" and people would be fooled even harder than they already are. There's got to be a fix for the underlying quality issues before replication becomes the next best thing to do.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#208

Earlier quoted context omitted.

> The challenge is there really isn't a good way to incentivize that work. What if we got Undergrads (with hope of graduate studies) to do it? Could be a great way to train them on the skills required for research without the pressure of it also being novel?

Unfortunately, that might just lead to a bunch of type II errors instead, if an effect requires very precise experimental conditions that undergrads lack the expertise for.

Could it be useful as a first line of defence? A failed initial reproduction would not be seen as disqualifying, but it would bring the paper to the attention of more senior people who could try to reproduce it themselves. (Maybe they still wouldn't bother, but hopefully they'd at least be more likely to.)

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#209

Earlier quoted context omitted.

Isn't disqualifying X months of potentially great research due to a misformed, but existing reference harsh? I don't think they'd be okay with references that are actually made up.

Science relies on trust.. a lot. So things which show dishonesty are penalised greatly. If we were to remove trust then peer reviewing a paper might take months of work or even years.

And that timeline only grows with the complexity of the field in question. I think this is inherently a function of the complexity of the study, and rather than harshly penalizing such shortcomings we should develop tools that address them and improve productivity. AI can speed up the verification of requirements like proper citations, both on the author's and reviewer's side.

Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers

#210
post #4

If these are so easy to identify, why not just incorporate some kind of screening into the early stages of peer review?

What makes you believe that are easy to identify?

Isn't that what GPTZero does?
Post reply on HN