Live data from Hacker News

Major AI conference flooded with peer reviews written by AI

nature.com

61–70 of 140 posts

Re: Major AI conference flooded with peer reviews written by AI

#61
post #24

> Pangram’s analysis revealed that around 21% of the ICLR peer reviews were fully AI-generated, and more than half contained signs of AI use. The findings were posted online by Pangram Labs. “People were suspicious, but they didn’t have any concrete proof,” says Spero. “Over the course of 12 hours, we wrote some code to parse out all of the text content from these paper submissions,” he adds. But what's the proof? Ho…

I wouldn't be surprised to learn that the AI detection tool is itself an AI

Fighting fire with fire sounds good in theory but in the end you're still on fire.

Re: Major AI conference flooded with peer reviews written by AI

#62

Earlier quoted context omitted.

The argument is that there is no incentive to carefully review a paper (I agree), however what used to occur is people would do the right thing without explicit incentives. This has totally disappeared.

The concept of the professional has been basically obliterated in our society. Instead we have people doing engineering, science, and doctoring as, just, jobs. Individual contributors of various flavors to be shuffled around by middle management. Without professions, there are no more professional communities really, no more professional standards to uphold, no reason to get in the way of somebody’s publications.

It is soundly unfair and unjustified to extrapolate the ML community to all professions. What is happening in the ML world is the exception, not the norm, and not some fundamental failing of society.

Re: Major AI conference flooded with peer reviews written by AI

#63
post #24

> Pangram’s analysis revealed that around 21% of the ICLR peer reviews were fully AI-generated, and more than half contained signs of AI use. The findings were posted online by Pangram Labs. “People were suspicious, but they didn’t have any concrete proof,” says Spero. “Over the course of 12 hours, we wrote some code to parse out all of the text content from these paper submissions,” he adds. But what's the proof? Ho…

"proof" was an unfortunate phrase to use. However, a proper statistical analysis can be objective. And these kinds of tools are perfectly suited to such an analysis.

Re: Major AI conference flooded with peer reviews written by AI

#64
This is also the conference where everybody was briefly deanonymized due to an OpenReview bug: https://eu.36kr.com/en/p/3572028126116993 Now all the review scores have been reset, and new area chairs will make all decisions from scratch based on the reviews and authors' responses.

Re: Major AI conference flooded with peer reviews written by AI

#65
I couldn't care less tbh. I just want to know whether they're correct or not. We need something like unit testing and integration testing, but for ideas.

For the record I actually like the AI writing style. It's a huge improvement in readability over most academic writing I used to come across.

Re: Major AI conference flooded with peer reviews written by AI

#66

This is the kind of situation where everything sucks. You'd think that one of the biggest AI conference out there would have seen this coming. On the one hand (and the most important thing, IMO) it's really bad to judge people on the basis of "AI detectors", especially when this can have an impact on their career. It's also used in education, and that sucks even more. AI detectors have bad rates, can't detect concent…

Not just spell checking, but translation. English is not the first language for most of the reviewers.

But you can see the slippery slope: first you ask your favorite LLM to check your grammar, and before you think about it, you are just asking it to write the whole thing.

Re: Major AI conference flooded with peer reviews written by AI

#67
post #6

While I think there's significant AI "offloading" in writing, the article's methodology relies on "AI-detectors," which reads like PR for Pangram. I don't need to explain why AI detectors are mostly bullshit and harmful for people who have never used LLMs. [1] 1: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

I am not sure if you are familiar with Pangram (co-founder here) but we are a group of research scientists who have made significant progress in this problem space. If your mental model of AI detectors is still GPTZero or the ones that say the declaration of independence is AI, then you probably haven't seen how much better they've gotten. This paper by economists from the University of Chicago economists found zero…

I thought the author was attempting to highlight the hypocrisy of using an AI to detect other uses of AI, as if one was a good use, and the other bad.

Re: Major AI conference flooded with peer reviews written by AI

#68

I wouldn't be surprised if the headline is accurate, but AI detectors are widely understood to be unreliable, and I see no evidence that this AI detector has overcome the well-deserved stigma.

Co-founder of Pangram here. Our false positive rate is typically around 1 in 10,000. https://www.pangram.com/blog/all-about-false-positives-in-ai... . We also wanted to quantify our EditLens model's FPR on the same domain, so we ran all of ICLR's 2022 reviews. Of 10,202 reviews, Pangram marked 10,190 as fully human, 10 as lightly AI-edited, 1 as moderately AI-edited, 1 as heavily AI-edited, and none as fully AI-gener…

Give your final sentence a re-read there....

Re: Major AI conference flooded with peer reviews written by AI

#69
post #59

Earlier quoted context omitted.

I am not sure if you are familiar with Pangram (co-founder here) but we are a group of research scientists who have made significant progress in this problem space. If your mental model of AI detectors is still GPTZero or the ones that say the declaration of independence is AI, then you probably haven't seen how much better they've gotten. This paper by economists from the University of Chicago economists found zero…

The response would be more helpful if it directly addresses the arguments in posts from that search result.

There are dozens of first generation AI detectors and they all suck. I'm not going to defend them. Most of them use perplexity based methods, which is a decent separators of AI and human text (80-90%) but has flaws that can't be overcome and high FPRs on ESL text.

https://www.pangram.com/blog/why-perplexity-and-burstiness-f...

Pangram is fundamentally different technology, it's a large deep learning based model that is trained on hundreds of millions of human and AI examples. Some people see a dozen failed attempts at a problem as proof that the problem is impossible, but I would like to remind you that basically every major and minor technology was preceded by failed attempts.

Re: Major AI conference flooded with peer reviews written by AI

#70
post #3

> Controversy has erupted after 21% of manuscript reviews for an international AI conference were found to be generated by artificial intelligence. 21%...? Am I reading it right? I bet no one expected it's so low when they clicked this title.

21% fully AI generated. In other words, 21% blatant fraud. In accident investigation we often refer to "holes in the swiss cheese lining up." Dereliction of duty is commonly one of the holes that lines up with all the others, and is apparently rampant in this field.

Who cares what tool was used to write the work? The important question is what percentage of reviews found errors or provided valuable feedback. The important metric is whether or not it did the job, not how it was produced.

I think there is a far more interesting discussion to be had here about how useful the 21% percent were. How well does an AI execute a peer review?

Post reply on HN