Earlier quoted context omitted.
The problem is that reproducing something is really, really hard! Even if something doesn't reproduce in one experiment, it might be due to slight changes in some variables we don't even think about. There are some ways to circumvent it (e.g. team that's being reproduced cooperating with reproducing team and agreeing on what variables are important for the experiemnt and which are not), but it's really hard. The solu…
Every time some easy "Reproducibility is hard / not worth the effort" I hear "The original research wasn't meaningful or valuable".
GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
311–320 of 528 posts
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#312Earlier quoted context omitted.
In my mental model, the fundamental problem of reproducibility is that scientists have very hard time to find a penny to fund such research. No one wants to grant “hey I need $1m and 2 years to validate the paper from last year which looks suspicious”. Until we can change how we fund science on the fundamental level; how we assign grants — it will be indeed very hard problem to deal with.
In theory, asking grad students and early career folks to run replications would be a great training tool. But the problem isn’t just funding, it’s time. Successfully running a replication doesn’t get you a publication to help your career.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#313Earlier quoted context omitted.
In my mental model, the fundamental problem of reproducibility is that scientists have very hard time to find a penny to fund such research. No one wants to grant “hey I need $1m and 2 years to validate the paper from last year which looks suspicious”. Until we can change how we fund science on the fundamental level; how we assign grants — it will be indeed very hard problem to deal with.
Funding is definitely a problem, but frankly reproduction is common. If you build off someone else's work (as is the norm) you need to reproduce first. But without repetition being impactful to your career and the pressure to quickly and constantly push new work, a failure to reproduce is generally considered a reason to move on and tackle a different domain. It takes longer to trace the failure and the bar is higher…
Not if the result you're building off of is a model, you can just assume it
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#314I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the pa…
>this error does make me pause to wonder how much of the rest of the paper used AI assistance And this is what's operative here. The error spotted, the entire class of error spotted, is easily checked/verified by a non-domain expert. These are the errors we can confirm readily, with obvious and unmistakable signature of hallucination. If these are the only errors, we are not troubled. However: we do not know if these…
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#315Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#316Earlier quoted context omitted.
>this error does make me pause to wonder how much of the rest of the paper used AI assistance And this is what's operative here. The error spotted, the entire class of error spotted, is easily checked/verified by a non-domain expert. These are the errors we can confirm readily, with obvious and unmistakable signature of hallucination. If these are the only errors, we are not troubled. However: we do not know if these…
This seems like finding spelling errors and using them to cast the entire paper into doubt. I am unconvinced that the particular error mentioned above is a hallucination, and even less convinced that it is a sign of some kind of rampant use of AI. I hope to find better examples later in the comment section.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#317Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#318The ironic part about these hallucinations is that a research paper includes a literature review because the goal of the research is to be in dialogue with prior work, to show a gap in the existing literature, and to further the knowledge that this prior work has built. By using an LLM to fabricate citations, authors are moving away from this noble pursuit of knowledge built on the "shoulders of giants" and show that…
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#319Earlier quoted context omitted.
In my mental model, the fundamental problem of reproducibility is that scientists have very hard time to find a penny to fund such research. No one wants to grant “hey I need $1m and 2 years to validate the paper from last year which looks suspicious”. Until we can change how we fund science on the fundamental level; how we assign grants — it will be indeed very hard problem to deal with.
I often think we should movefrom peer review as "certification" to peer review as "triage", with replication determining how much trust and downstream weight a result earns over time.
academia is too fragmented and extremely inefficient
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#320Yuck, this is going to really harm scientific research. There is already a problem with papers falsifying data/samples/etc, LLMs being able to put out plausible papers is just going to make it worse. On the bright side, maybe this will get the scientific community and science journalists to finally take reproducibility more seriously. I'd love to see future reporting that instead of saying "Research finds amazing che…
In my mental model, the fundamental problem of reproducibility is that scientists have very hard time to find a penny to fund such research. No one wants to grant “hey I need $1m and 2 years to validate the paper from last year which looks suspicious”. Until we can change how we fund science on the fundamental level; how we assign grants — it will be indeed very hard problem to deal with.
of course the problem is that academia likes to assert its autonomy (and grant orgs are staffed by academia largely)