Which is worse: a) p-hacking and suppressing null results b) hallucinations c) falsifying data Would be cool to see an analysis of this
GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
181–190 of 528 posts
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#182Getting papers published is now more about embellishing your CV versus a sincere desire to present new research. I see this everywhere at every level. Getting a paper published anywhere is a checkbox in completing your resume. As an industry we need to stop taking this into consideration when reviewing candidates or deciding pay. In some sense it has become an anti-signal.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#183This suggests that nobody was screening this papers in the first place—so is it actually significant that people are using LLMs in a setting without meaningful oversight? These clearly aren't being peer-reviewed, so there's no natural check on LLM usage (which is different than what we see in work published in journals).
As one who reviews 20+ papers per year, we don't have time to verify each reference. We verify: is the stuff correct, and is it worthy of publication (in the given venue) given that it is correct. There is still some trust in the authors to not submit made-up-stuff, albeit it is diminishing.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#184Earlier quoted context omitted.
> I'd love to see future reporting that instead of saying "Research finds amazing chemical x which does y" you see "Researcher reproduces amazing results for chemical x which does y. First discovered by z". Most people (that I talk to, at least) in science agree that there's a reproducibility crisis. The challenge is there really isn't a good way to incentivize that work. Fundamentally (unless you're independent weal…
> Eventually you may arrive at something like the H-index, which is defined as "The highest number H you can pick, where H is the number of papers you have written with H citations." It's the Google search algorithm all over again. And it's the certificate trust hierarchy all over again. We keep working on the same problems. Like the two cases I mentioned, this is a matter of making adjustments until you have the des…
First X people that reproduce Y get Z percent of patent revenue.
Or something similar.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#185Earlier quoted context omitted.
In my mental model, the fundamental problem of reproducibility is that scientists have very hard time to find a penny to fund such research. No one wants to grant “hey I need $1m and 2 years to validate the paper from last year which looks suspicious”. Until we can change how we fund science on the fundamental level; how we assign grants — it will be indeed very hard problem to deal with.
In theory, asking grad students and early career folks to run replications would be a great training tool. But the problem isn’t just funding, it’s time. Successfully running a replication doesn’t get you a publication to help your career.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#186Earlier quoted context omitted.
> the content of the papers themselves are not necessarily invalidated. For example, authors may have given an LLM a partial description of a citation and asked the LLM to produce bibtex (a formatted reference) Maybe I'm overreacting, but this feels like an insanely biased response. They found the one potentially innocuous reason and latched onto that as a way to hand-wave the entire problem away. Science already had…
Isn't disqualifying X months of potentially great research due to a misformed, but existing reference harsh? I don't think they'd be okay with references that are actually made up.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#187Earlier quoted context omitted.
> The challenge is there really isn't a good way to incentivize that work. Ban publication of any research that hasn't been reproduced.
> Ban publication of any research that hasn't been reproduced. Unless it is published, nobody will know about it and thus nobody will try to reproduce it.
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#188Earlier quoted context omitted.
Why not run every submitted paper through GPTZero (before sending to reviewers) and summarily reject any paper with a hallucination?
That's how GPTZero wants to situate themselves. Who would pay them? Conference organizers are already unpaid and undestaffed, and most conferences aren't profitable. I think rejections shouldn't be automatic. Sometimes there are just typos. Sometimes authors don't understand BibTeX. This needs to be done in a way that reduces the workload for reviewers. One way of doing this would be for GPTZero to annotate each pape…
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#189Yuck, this is going to really harm scientific research. There is already a problem with papers falsifying data/samples/etc, LLMs being able to put out plausible papers is just going to make it worse. On the bright side, maybe this will get the scientific community and science journalists to finally take reproducibility more seriously. I'd love to see future reporting that instead of saying "Research finds amazing che…
Re: GPTZero finds 100 new hallucinations in NeurIPS 2025 accepted papers
#190Earlier quoted context omitted.
As one who reviews 20+ papers per year, we don't have time to verify each reference. We verify: is the stuff correct, and is it worthy of publication (in the given venue) given that it is correct. There is still some trust in the authors to not submit made-up-stuff, albeit it is diminishing.
Sorry, but if someone makes a claim and cites a reference, how do you verify "is the stuff correct" without checking that reference?
Fake references are more common in the introduction where you list relevant material to strengthen your results. They often don't change the validity of the claim, but the potential impact or value.