Live data from Hacker News

New arXiv policy: 1-year ban for hallucinated references

twitter.com

81–90 of 242 posts

Re: New arXiv policy: 1-year ban for hallucinated references

#81

Earlier quoted context omitted.

You could at least filter out hallucinated references which simply don't exist pretty trivially, I'd imagine.

It's more than that. if there are mistakes, then you can also be flagged. read the whole tweet: If generative AI tools generate inappropriate language, plagiarized content, biased content, errors, mistakes, incorrect references, or misleading content, and that output is included in scientific works, it is the responsibility of the author(s).

If you'd read the whole series of tweets it's obvious that is not their intention and there needs to be "incontrovertible evidence that the authors did not check the results of LLM generation" for the penalty to apply.

It's not hard to divine their intentions: you are entirely responsible for what you summit and if it's clearly slop(py) you get a ban. In a reply they state that they are seeking to apply this rule fairly and accurately and are mindful of unintended effects.

Re: New arXiv policy: 1-year ban for hallucinated references

#83
post #21

Earlier quoted context omitted.

It's not the kind of mistake that is possible unless you're engaging in fraud anyway.

> It's not the kind of mistake that is possible unless you're engaging in fraud anyway. Seriously? You can't fathom an honest researcher asking for AI to find a citation they know exists, and the AI inserting or modifying a citation incorrectly without them realizing? If you find evidence of fraud by all means lay down the hammer. Using a single hallucinated citation like it's some kind of ironclad proxy just because…

if an llm does the work, you did not write it or research it, the llm did. you have no business crediting yourself as an author.

if someone writes a paper and an entirely different person takes credit for it without even bothering to check if the actual writer just made shit up, they deserve a lifetime ban. seems like a year is a very light punishment.

Re: New arXiv policy: 1-year ban for hallucinated references

#85
post #33

Earlier quoted context omitted.

> It's just one darn hallucinated citation for heaven's sake, not fraud or something. It is fraud. > It doesn't account for the substance or quality of their work at all. References are part of the work. If you're making up the references, what else are you making up? > People make mistakes and a good fraction of them can learn from those mistakes. There's no need to permanently cripple someone's ability to progress…

it's very silly, but not a big deal. Arxiv is becoming irrelevant these days anyways. In fact would be better if they just banned AI, so we could just get off the luddite platforms. Automated research is the future, end of story. And really it couldn't have come out at a better time, given the increasingly diminishing returns on human powered research.

If automated research is the future, it has to be research, not making stuff up.

Which of those two does "hallucinated references" fit into?

Re: New arXiv policy: 1-year ban for hallucinated references

#86
post #22

Earlier quoted context omitted.

A "mistake" would be a typo in a real citation. A hallucinated citation is evidence of just plain laziness and negligence, which taints the entire submission.

No it is not. Seriously. All you need for this to happen is for your lab partner to ask AI to add a missing citation that they are already familiar with at the last minute before a midnight submission deadline, and for the AI to hallucinate something else, and for them to honestly miss this. It does not even imply any involvement on your part, let alone that either of you were lazy or negligent on the actual research…

You’re confusing the issue here by saying it’s not your fault, it’s your lab partner’s. We’re talking about why your lab partner did something wrong. You can assign blame for the wrong thing separately.

The citation is part of the substance of the paper. If you YOLOed in a citation without checking it, seems justified to suspect that you may have YOLOed in some data, or some analysis, or maybe even the conclusion.

Re: New arXiv policy: 1-year ban for hallucinated references

#87
post #36

Earlier quoted context omitted.

No it is not. Seriously. All you need for this to happen is for your lab partner to ask AI to add a missing citation that they are already familiar with at the last minute before a midnight submission deadline, and for the AI to hallucinate something else, and for them to honestly miss this. It does not even imply any involvement on your part, let alone that either of you were lazy or negligent on the actual research…

There are no deadlines for journal submissions. Even if you felt you were running close to your revisions being due, an email to an editor will probably fix this for you. And what you described is still negligent, not verifying the garbage output bot did not in fact output garbage.

Even more, there are no deadlines for arXiv submissions.

Re: New arXiv policy: 1-year ban for hallucinated references

#88
post #67

Good. If it’s not worth your time to check the output of your LLM carefully, it’s not worth my time to read it.

Unfortunately, it's probably not worth your time to read 99% of arxiv papers, LLM generated or otherwise.

Ever pick a random one and really dive in?

Re: New arXiv policy: 1-year ban for hallucinated references

#89
post #5

It seems a good idea to ban cheating, but how hard is it, especially in new reasoning/agents contexts to validate references? The deeper question is whether legitimate AI generated results are allowed or not? Test - In the extreme - think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not?

In that case, you would just not do a reference. End to end autonomous science might have fewer concrete citations as the contributing knowledge is just the sum of the training data of the model.

Re: New arXiv policy: 1-year ban for hallucinated references

#90

Earlier quoted context omitted.

Fraud requires intent to deceive _or_ reckless disregard, sometimes called, “conscious indifference” for the veracity of the statement asserted.

No. One single hallucinated citation on a document with you as an author is not evidence of your reckless disregard for anything. These exaggerations are crazy and you would absolutely deny such accusations if you missed your co-author's AI hallucinating a citation on your manuscript too. At best it would be careless , if you really relish extrapolating from one data point and smearing people's character based on tha…

I’ve disagreed with some of your other stances in this thread, but I want to acknowledge the validity of your take here.

You’re right that a single hallucinated line is not evidence of reckless disregard - because that could have happened on a final follow-up pass after you had performed due diligence. It’s happened to me. I know how challenging it can be to keep bad patterns out of LLM generated output, because human communication is full of bad patterns. It’s a constant battle, and sometimes I suspect that my hard-line posture actually encourages the LLM to regularly “vibe check” me! E.g. “Are you sure you’re really the guy you’re trying to be? Because if you are you wouldn’t miss this.” LLMs are devious, and that’s why I respect them so much. If you think they’re pumping the breaks then you should check again, because they probably just put the pedal to the metal.

That being said, I regularly insist on doing certain things myself. If I were publishing a paper intended to be taken seriously - citations would be one of the things I checked manually. But I can easily see myself doing a final follow-up pass after everything looks perfect, and missing a last minute change. I would hope that I would catch that, but when you’re approaching the finish line - that’s when you expect your team to come together. That’s when everything is “supposed to” fall into place. It’s the last place you would expect to be sabotaged, and in hindsight, probably the best place to be a saboteur.

Post reply on HN