Live data from Hacker News

New arXiv policy: 1-year ban for hallucinated references

twitter.com

121–130 of 242 posts

Re: New arXiv policy: 1-year ban for hallucinated references

#121
post #109

Earlier quoted context omitted.

Fraud in the scientific world has generally taken the form of fabricated results, but I don't agree that the word has transitioned away from the common and legal meaning of deception in order to get a benefit. Even if it had though, I'd be perfectly comfortable calling this fraud in this discussion based on the common meaning of the word. Just because we're talking about a scientific context does not mean we need to…

I disagree with your implicit assertion regarding the common meaning of the term in this context. I believe that the term fraud as commonly used when discussing things in a scientific context has always (for at least my entire life) been taken to refer to knowingly and intentionally falsified research results (also falsified appointments, falsified affiliations, falsified authorships, etc). > deception in order to ge…

We're not "discussing things in a scientific context" here. We're in the context of a startup/programmer news aggregator discussing scientific news. We are not "a bunch of tenured career researchers" discussing amongst ourselves so the jargon appropriate for that context is not the appropriate jargon - rather we need to use the jargon that the startup/programmer news aggregator crowd would understand.

That said even in a scientific context I still disagree and your example at the end is a fine starting point. By comparison imagine one of the profs said told the others that their house was burgled. The others would probably be thinking that things like TVs or computers or money was stolen, and not that the thief simply stole all their spoons. That doesn't make having all your spoons stolen not burglary. Likewise the profs expect that the results or authorships are where the fraud occurred because those are the best places to extract value with fraud, not by avoiding the simple act of writing the paper with correct citations. That doesn't mean fraudulently using an LLM to hallucinate a paper from your (we'll suppose for sake of argument) actual results is any less fraudulent though, it's just an unexpected form of fraud.

Edit: I want to be clear that this is not my argument: "well he exhibited reckless disregard for his professional duties when he opted not to bother reading the citation section". I see other people making that argument, and I'm not sure if they're right or wrong that that's another reason why it is fraud, but I'm certain that we don't even need to reach that question.

My argument is that it is fraud to represent the paper as a scholarly work when you don't know that it is correct. It is not that you are taking a risk it might be wrong, it is that you are actively representing that you know it is correct and if you do not know that you are committing fraud even if it happens to be so. This is a case of intentional deception, the deception being the representation that this is scholarly work, not reckless disregard for the truth as to the accuracy of the citations.

Re: New arXiv policy: 1-year ban for hallucinated references

#122
post #48

Earlier quoted context omitted.

> It's not the kind of mistake that is possible unless you're engaging in fraud anyway. Seriously? You can't fathom an honest researcher asking for AI to find a citation they know exists, and the AI inserting or modifying a citation incorrectly without them realizing? If you find evidence of fraud by all means lay down the hammer. Using a single hallucinated citation like it's some kind of ironclad proxy just because…

If you are citing a work you paste a citation to that work. If you are bullshitting you ask an AI to come up with a citation. Jesus, there is zero reason to ever "generate a citation" if you are not, in fact, commiting fraud.

That's like saying that there's zero reason to ever ask an LLM to do basic math for you. Sure you probably shouldn't do that but sometimes it's convenient and so people will inevitably do exactly that regardless of the somewhat frequent wrong answers that are guaranteed to ensue.

Re: New arXiv policy: 1-year ban for hallucinated references

#123

how will they detect hallucinated refs at scale? Manual spot checks? Automated DOI verification? The policy seems right but enforcement is the hard part.

However difficult it might be right now it's only going to get easier. Anyway I don't think proactive enforcement is the point. Rather now they have an official method by which to address incidents that are brought to their attention.

Re: New arXiv policy: 1-year ban for hallucinated references

#124
post #4

> The penalty is a 1-year ban from arXiv followed by the requirement that subsequent arXiv submissions must first be accepted at a reputable peer-reviewed venue. This is incredibly good for science. arXiv is free, but it's a privilege not a right! I'm not seeing this clearly listed on https://info.arxiv.org/help/policies/index.html so it's possible this is planned but not live yet - or perhaps I'm not digging deeply…

My take: this seems excessive. ArXiv doesn't even check the submission closely, so how can they know? They say "errors, mistakes" They use an automated system to check if the basic requirements were met, and sometimes papers are flagged for further superficial human review, but there is no way they can possibly do this at scale or check every reference. This would be like trying to do peer review, but for a preprint…

This puts the burden to make sure it's right on the submitter, where it should be. Verification can come at any time after that; the submitter understands the consequences of hallucinated references. Verification can be crowd-sourced (and likely will be).

Nothing stops someone from putting a PDF on the internet. I'm fine with ArXiv holding a high standard.

Re: New arXiv policy: 1-year ban for hallucinated references

#125
post #52

Earlier quoted context omitted.

> No, it is emphatically not. D Fraud requires intent to deceive. I'm about as pro AI-as-a-research--and-writing-assistant and anti AI-witchhunt as they come, but I simply cannot parse what I've quoted here. Posting slop to arxiv is blatant deception. Posting an article is an attestation that the article is a genuine engagement with the literature . If you're posting things to arxiv that are not sincere engagements w…

>I'm about as pro AI-as-a-research--and-writing-assistant and anti AI-witchhunt as they come, but I simply cannot parse what I've quoted here. Ditto. And its only 1 year. Like its about the most reasonable thing they could have done.

> And its only 1 year

No, it emphatically is not just a year! It's perpetual, and that's literally been my entire point this whole time. If it was just one year I would've had no complaints - and I made that clear from the very first comment!

What part of "...followed by the requirement that subsequent arXiv submissions must first be accepted at a reputable peer-reviewed venue..." is everyone here reading and still somehow interpreting to be limited to 1 year?

Re: New arXiv policy: 1-year ban for hallucinated references

#126
post #33

Earlier quoted context omitted.

> This is incredibly good for science. I disagree. It's just one darn hallucinated citation for heaven's sake, not fraud or something. It doesn't account for the substance or quality of their work at all. A one-year ban seems plenty sufficient for a minor first time mistake like this. People make mistakes and a good fraction of them can learn from those mistakes. There's no need to permanently cripple someone's abili…

> It's just one darn hallucinated citation for heaven's sake, not fraud or something. It is fraud. > It doesn't account for the substance or quality of their work at all. References are part of the work. If you're making up the references, what else are you making up? > People make mistakes and a good fraction of them can learn from those mistakes. There's no need to permanently cripple someone's ability to progress…

If you write your own paper (mostly) and choose your own references (because you've actually read the papers) you won't have a problem.

Re: New arXiv policy: 1-year ban for hallucinated references

#127
post #118
post #96

Earlier quoted context omitted.

I bet, since this has been posted, someone here has already vibe coded a reference checker that they plan to put behind a subscription. This is good for reference checking, but I doubt this will do much for the most likely shoddy science that accompanies hallucinated references.

The frontier LLMs are getting pretty good at checking this sort of thing. You could prompt them to not only verify the references are real but that they actually state what the article claims. Some human review will still be needed but I'll bet this approach could find a lot of academic fraud.

> The frontier LLMs are getting pretty good at checking this sort of thing.

No, this is career ending high stakes. it requires old school "actually check a record of reality" type methods, like a database query or http get to one of the many services that hold this info.

Re: New arXiv policy: 1-year ban for hallucinated references

#128
post #63

Earlier quoted context omitted.

No. One single hallucinated citation on a document with you as an author is not evidence of your reckless disregard for anything. These exaggerations are crazy and you would absolutely deny such accusations if you missed your co-author's AI hallucinating a citation on your manuscript too. At best it would be careless , if you really relish extrapolating from one data point and smearing people's character based on tha…

If your co-author inserted the fradulent reference, I agree that you may not have committed fraud. But your co-author did, and you didn't check their work. and knowing that you didn't check their work, you signed off on it. You didn't pick your co-author very well, but arXiv lacks investigative powers to determine which co-author did the bad, so they all get the consequence.

Do you think every co-author on a 100-author paper checks every citation? It's like saying that every member of a large software team personally reviews every line of code. It's just completely divorced from reality.

Re: New arXiv policy: 1-year ban for hallucinated references

#129

Earlier quoted context omitted.

I’ve disagreed with some of your other stances in this thread, but I want to acknowledge the validity of your take here. You’re right that a single hallucinated line is not evidence of reckless disregard - because that could have happened on a final follow-up pass after you had performed due diligence. It’s happened to me. I know how challenging it can be to keep bad patterns out of LLM generated output, because huma…

You're saying it as if the poor author just had no choice but to let LLM write their bibliography. To avoid hallucinations, maybe just don't let an LLM write any part of your paper? You can only get in this situation if you let a bullshit generator write your paper, and the fraud is that you are generating bullshit and calling it a paper. No buts. It's impossible to trigger this accidentally, or without reckless disr…

Calling LLMs "bullshit generators" in the year 2026 just shows a lack of seriousness.

Re: New arXiv policy: 1-year ban for hallucinated references

#130

Earlier quoted context omitted.

No it is not. Seriously. All you need for this to happen is for your lab partner to ask AI to add a missing citation that they are already familiar with at the last minute before a midnight submission deadline, and for the AI to hallucinate something else, and for them to honestly miss this. It does not even imply any involvement on your part, let alone that either of you were lazy or negligent on the actual research…

In other words: all it needs for your paper to have fraud is for your lab partner to add fraud to your paper. I'm not seeing the problem here. The only problem is that your lab partner should be banned and not you. But being incentivised to check your co-author's work before submission isn't a bad thing.

> But being incentivised to check your co-author's work before submission isn't a bad thing.

Nobody was arguing against this.

Post reply on HN