Live data from Hacker News

2% of ICML papers desk rejected because the authors used LLM in their reviews

blog.icml.cc

91–100 of 172 posts

Re: 2% of ICML papers desk rejected because the authors used LLM in their reviews

#91
post #35

Earlier quoted context omitted.

> Interesting, so someone submitting a paper for review could also submit one with hidden instructions for LLMs to summarise or review it in a very positive light. I may or may not know a guy who added several hidden sentences in Finnish to his CV that might have helped him in landing an interview.

>several hidden sentences in Finnish Is this a reference to something?

Not at all. It's just that reportedly LLMs used to have a blind spot for prompt injection in languages with relatively few speakers and grammar dissimilar to that of English.

Re: 2% of ICML papers desk rejected because the authors used LLM in their reviews

#92

One thing to note. They were quite conservative in their approach, so the only things that were rejected were from people who had agreed not to use an LLM and almost definitely did use an LLM (since they fed hidden watermarked instructions to the llm's). This means the true number of people that used LLM's in their review (even in group A that had agreed not to) is likely higher. Also worth noting, 10% of these autho…

Yes for those in group B I'd suspect many were doing exactly what these cheaters in group A were doing - submitting the unaltered output of an LLM as their review.

The rejection is based on the dishonesty of explicitly committing to standard A and then knowingly violating it, not on LLM use as such. I think that's pretty fair, considering that everyone could have just chosen B if they wanted to.

Re: 2% of ICML papers desk rejected because the authors used LLM in their reviews

#94

Earlier quoted context omitted.

It's not a fully consensus view, but a majority of sociologists agree that high severity deterrence has limited effectiveness against crime. Instead, certainty of enforcement is the most salient factor.

Deterrence is only part of it. It's morally instructive, it tells people that they live in a society that takes rules seriously.

What is the aim of "moral instruction" if not deterrence? Surely it needs be instruction in pursuit of an outcome?

Re: 2% of ICML papers desk rejected because the authors used LLM in their reviews

#95

Earlier quoted context omitted.

I was thinking this too, but I don't believe this is the case, and I feel like it would not be a good idea either. Most of these people are likely students; this should be a learning moment, but I don't think it is yet grounds for their entire academic career to be crippled by being unable to publish in a top-tier ML venue.

If this is tolerated, it sends exactly the wrong kind of message. The students, if they are, should be banned for life. Let them serve as an example for myriads of future students, this will be a better outcome in the long run. This didn't trip for people who were merely bouncing ideas off a LLM, they caught people who copy and pasted straight from their LLM.

Well, maybe they found themselves in the last hours of the deadline without the reviews done... in some cases due to procrastination, but in a few cases perhaps because life is hard and they just couldn't do it. So they used the LLM as a last resort to not go beyond deadline (which I assume maybe was penalized as well?)

To err is human, it makes sense that they are punished (and the harshest part of the punishment is not having a paper rejected, it's the loss of face with coauthors and others, BTW. Face is important in academia) but "for life" is way too much IMO.

Re: 2% of ICML papers desk rejected because the authors used LLM in their reviews

#96
It would be interesting to know how many of the cheaters didn't check policy A, but checked "don't care if A or B". Because the operative part of that is "don't care", not "I will strictly adhere to either policy A or B, whatever somebody else selects for me".

So it is a sneaky and typically academic way of doing stuff. Also, "We hope that by taking strong action against violations of agreed-upon policy we will remind the community that as our field changes rapidly the thing we must protect most actively is our trust in each other. If we cannot adapt our systems in a setting based in trust, we will find that they soon become outdated and meaningless." is so academic and pointless.

Re: 2% of ICML papers desk rejected because the authors used LLM in their reviews

#97
post #89

Earlier quoted context omitted.

I'm not sure what experience anyone in this thread has with grad level research as a student/author, but I can assure you that heads roll over this kind of thing. A professor's career is built on reputation, and that reputation is as strong as their students' (who do much of the "work" such as it is). It comes down to the professor, but this can be a career-ending moment for those students and I'm quite confident the…

[flagged]

This comment doesn't seem to fit the discussion at all?

The discussion is not about humans using LLLs to write papers. It is about humans who agreed not to use LLVM in reviewing papers, then did exactly that.

Re: 2% of ICML papers desk rejected because the authors used LLM in their reviews

#98
post #89

Earlier quoted context omitted.

I'm not sure what experience anyone in this thread has with grad level research as a student/author, but I can assure you that heads roll over this kind of thing. A professor's career is built on reputation, and that reputation is as strong as their students' (who do much of the "work" such as it is). It comes down to the professor, but this can be a career-ending moment for those students and I'm quite confident the…

[flagged]

They agreed to the no LLM policy.

Re: 2% of ICML papers desk rejected because the authors used LLM in their reviews

#99
post #47

Great experiment! Correct me if I'm wrong, but this means that many people are using LLMs despite claiming not to. It's the first symptom of a dependency mechanism. If this happens in this context, who knows what happens in normal work or school environments? (P.S.: The use of watermarks in PDFs to detect LLM usage is very interesting, even though the LLM might ignore hidden instructions.)

The rate was 1% so this does not mean that "many" people are using LLMs despite claiming not to.

Re: 2% of ICML papers desk rejected because the authors used LLM in their reviews

#100
This is about reviewers, not authors. Title is a bit misleading.

In any case, having reviewed a lot of mostly very poorly written articles and occasionally solid papers when I was still a researcher, I can sympathize with using LLMs to streamline the process. There are a lot of meh papers that are OK for a low profile workshop or small conference where you cut people some slack. But generally standards should be higher for things like journals. Judging what is acceptable for what is part of the game. For a workshop, the goal is to get interesting junior researchers together with their senior peers. Honestly, workshops are where the action is in the academic world. You meet interesting people and share great ideas.

Most people may not realize this but there are a lot of people that are starting in their research career that will try to get their papers accepted for workshops, conferences, or journals. We all have to start somewhere. I certainly was not an amazing author early on. Getting rejections with constructive feedback is part of how you get better. Constructive feedback is the hard part of reviewing.

The more you publish, the more you get invited to review. It's how the process works. It generates a lot of work for reviewers. I reviewed probably at least 5-10 papers per month. It actually makes you a better author if you take that work seriously. But it can be a lot of work unless you get organized. That's on top of articles I chose to read for my own work. Digesting lots of papers efficiently is a key skill to learn.

Reviewing the good papers is actually relatively easy. It's enjoyable even; you learn something and you get to appreciate the amazing work the authors did. And then you write down your findings.

It's the mediocre ones that need a lot of careful work. You have to be fair and you have to be strict and right. And then you have to provide constructive feedback. With some journals, even an accept with revisions might land an article on the reject pile.

The bad ones are a chore. They are not enjoyable to read at all.

The flip side of LLMs is that both sides can and should (IMHO) use them: authors can use them to increase the quality of their papers. With LLMs there no longer is any excuse for papers with lots bad grammar/spelling or structure issues anymore. That actually makes review work harder. Because most submitted papers now look fairly decent which means you have to dive into the detail. Rejecting a very rough draft is easy. Rejecting a polished but flawed paper is not.

If I was still doing reviews (I'm not), I'd definitely use LLMs to pick apart papers, to quickly zoom in on the core issues and to help me keep my review fair and balanced and professional in tone. I would manually verify the most important bits and my effort would be proportional to which way I'm leaning based on what I know. Of course, editors can use LLMs as well to make sure reviews are fair and reasonable in their level of detail and argumentation. Reviewing the reviewers always has been a weakness of the peer review system and sometimes turf wars are being fought by some academics via reviews. It's one of the downsides of anonymous reviews and the academic world can be very political. A good editor would stay on top of this and deal with it appropriately.

LLMs are good at filtering, summarizing, flagging, etc. With proper guard rails, there's no reason to not lean on that a bit. It's the abuse that needs to be countered. In the end, that begins and ends with editors. They select the reviewers. So when those do a bad job, they need to act. And when their journals fill up with AI slop, it's their reputations that are on the line.

Like any tool, use caution and common sense. Blanket bans are not that productive at this stage.

Post reply on HN