Live data from Hacker News

New arXiv policy: 1-year ban for hallucinated references

twitter.com

11–20 of 242 posts

Re: New arXiv policy: 1-year ban for hallucinated references

#11
post #5

It seems a good idea to ban cheating, but how hard is it, especially in new reasoning/agents contexts to validate references? The deeper question is whether legitimate AI generated results are allowed or not? Test - In the extreme - think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not?

> think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not?

Sorry to be rude, but this seems like a dumb question. I want science to progress. A primary purpose of these journals is to progress science. A full proof of the Riemann Hypothesis progresses science. I don't care how it was produced, if Hitler is coauthor, etc, I just care that it is correct. Whether the authors should be rewarded for whatever methods they used can be a separate question.

Re: New arXiv policy: 1-year ban for hallucinated references

#12
post #5

It seems a good idea to ban cheating, but how hard is it, especially in new reasoning/agents contexts to validate references? The deeper question is whether legitimate AI generated results are allowed or not? Test - In the extreme - think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not?

If you use AI correctly, nobody should be able to tell that it was used at all.

Re: New arXiv policy: 1-year ban for hallucinated references

#13
post #5

It seems a good idea to ban cheating, but how hard is it, especially in new reasoning/agents contexts to validate references? The deeper question is whether legitimate AI generated results are allowed or not? Test - In the extreme - think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not?

[deleted]

Re: New arXiv policy: 1-year ban for hallucinated references

#14
post #5

It seems a good idea to ban cheating, but how hard is it, especially in new reasoning/agents contexts to validate references? The deeper question is whether legitimate AI generated results are allowed or not? Test - In the extreme - think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not?

> think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not? Sorry to be rude, but this seems like a dumb question. I want science to progress. A primary purpose of these journals is to progress science. A full proof of the Riemann Hypothesis progresses science. I don't care how it was produced, if Hitler is coauthor, etc, I just care that it is correct. Whether the…

Terence Tao had a nice talk from the Future of Mathematics conference posted yesterday [0] that shapes a lot of my own feelings on this matter.

The short of it is he argues how first to correctness shouldn't be the only goal / isn't a great optimisation incentive. Presentation and digestibility of correct results is a missing 1/3 when you've finished generation and verification. I completely agree with him. You don't just need an AI generated proof of the Reimann Hypothesis. You would really like it to be intentional and structured for others to understand.

A really beautiful quote I learned of in the talk is this:

> "We are not trying to meet some abstract production quota of definitions, theorems, and proofs. The measure of our success is whether what we do enables people to understand and think more clearly and effectively about math." - William Thurston

[0] https://www.youtube.com/watch?v=Uc2zt198U_U

Re: New arXiv policy: 1-year ban for hallucinated references

#16
post #6

Good; academic literature is in crisis because of all of the slop. Forcing some consequences on easily-detectable hallucinations can only be a good thing

It's not just AI, though. I did a doctorate in physics about 40 years back, and bad references were a problem back then.

In what way? Surely something like the source not quite saying what was cited, or mixing up citations, rather than inventing them outright?

Re: New arXiv policy: 1-year ban for hallucinated references

#17
post #5

It seems a good idea to ban cheating, but how hard is it, especially in new reasoning/agents contexts to validate references? The deeper question is whether legitimate AI generated results are allowed or not? Test - In the extreme - think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not?

There already exists multiple tools for automatically verifying references. This measure will likely only filter out the laziest and most incompetent of AI slop submissions. It's a very modest raising of the bar, but comes at zero cost to honest researchers.

I expect arXiv will still have problems with slop submissions but, at least, their references should actually exist going forward.

Re: New arXiv policy: 1-year ban for hallucinated references

#18
post #4

> The penalty is a 1-year ban from arXiv followed by the requirement that subsequent arXiv submissions must first be accepted at a reputable peer-reviewed venue. This is incredibly good for science. arXiv is free, but it's a privilege not a right! I'm not seeing this clearly listed on https://info.arxiv.org/help/policies/index.html so it's possible this is planned but not live yet - or perhaps I'm not digging deeply…

> This is incredibly good for science.

I disagree. It's just one darn hallucinated citation for heaven's sake, not fraud or something. It doesn't account for the substance or quality of their work at all. A one-year ban seems plenty sufficient for a minor first time mistake like this. People make mistakes and a good fraction of them can learn from those mistakes. There's no need to permanently cripple someone's ability to progress their life or contribute to humanity just because an AI hallucinated a reference one time in their life. That's punitive instead of rehabilitative.

Re: New arXiv policy: 1-year ban for hallucinated references

#19
post #5

It seems a good idea to ban cheating, but how hard is it, especially in new reasoning/agents contexts to validate references? The deeper question is whether legitimate AI generated results are allowed or not? Test - In the extreme - think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not?

It isn't "cheating" they're concerned with, it's sloppiness. This dictum isn't some sort of AI ban, but instead simply that if there is evidence that it was so low effort that the work includes such blatant problems, it's just adding noise.

Re: New arXiv policy: 1-year ban for hallucinated references

#20
post #6

Good; academic literature is in crisis because of all of the slop. Forcing some consequences on easily-detectable hallucinations can only be a good thing

It's not just AI, though. I did a doctorate in physics about 40 years back, and bad references were a problem back then.

Yes and ffs arrows kill people too but we don't bring that up every time we talk about what to do with guns.
Post reply on HN