Live data from Hacker News

Major AI conference flooded with peer reviews written by AI

nature.com

131–140 of 140 posts

Re: Major AI conference flooded with peer reviews written by AI

#131
post #53

Earlier quoted context omitted.

Nothing points out that the benchmark is invalid like a zero false positive rate. Seemingly it is pre-2020 text vs a few models rework of texts. I can see this model fall apart in many real world scenarios. Yes, LLMs use strange language if left to their own devices and this can surely be detected. 0% false positive rate under all circumstances? Implausible.

> Nothing points out that the benchmark is invalid like a zero false positive rate You’re punishing them for claiming to do a good job. If they truly are doing a bad job, surely there is a better criticism you could provide.

No one is punishing anyone. They just make an implausible claim. That is it.

Re: Major AI conference flooded with peer reviews written by AI

#132
post #6

While I think there's significant AI "offloading" in writing, the article's methodology relies on "AI-detectors," which reads like PR for Pangram. I don't need to explain why AI detectors are mostly bullshit and harmful for people who have never used LLMs. [1] 1: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

I am not sure if you are familiar with Pangram (co-founder here) but we are a group of research scientists who have made significant progress in this problem space. If your mental model of AI detectors is still GPTZero or the ones that say the declaration of independence is AI, then you probably haven't seen how much better they've gotten. This paper by economists from the University of Chicago economists found zero…

Hi Max! Thank you for updating my mental model of AI detectors.

I was with total certainty under the impression that detecting AI-written text to be an impossible-to-solve problem. I think that's because it's just so deceptively intuitive to believe that "for every detector, there'll just be a better LLM and it'll never stop."

I had recently published a macOS app called Pudding to help humans prove they wrote a text mainly under the assumption that this problem can't be solved with measurable certainty and traditional methods.

Now I'm of course a bit sad that the problem (and hence my solution) can be solved much more directly. But, hey, I fell in love with the problem, so I'm super impressed with what y'all are accomplishing at and with Pangram!

Re: Major AI conference flooded with peer reviews written by AI

#133
post #24

> Pangram’s analysis revealed that around 21% of the ICLR peer reviews were fully AI-generated, and more than half contained signs of AI use. The findings were posted online by Pangram Labs. “People were suspicious, but they didn’t have any concrete proof,” says Spero. “Over the course of 12 hours, we wrote some code to parse out all of the text content from these paper submissions,” he adds. But what's the proof? Ho…

You don’t. It’s bullshit inception.

Re: Major AI conference flooded with peer reviews written by AI

#134
post #92

Earlier quoted context omitted.

Yeah, Pangram does not provide any concrete proof, but it confirms many people's suspicions about their reviews. But it does flag reviews for a human to take a closer look and see if the review is flawed, low-effort, or contains major hallucinations.

> does not provide any concrete proof, but it confirms many people's suspicions Without proof there is no confirmation.

Formally? Sure. In the current zeitgeist it’s more than enough to start pointing fingers around, etc.

Re: Major AI conference flooded with peer reviews written by AI

#135

Earlier quoted context omitted.

That's a strategy for producing maximally convincing BS content, but the scientific method was absent. I'm sure it was an oversight... ; )

That’s on you. You get to decide what “best” means when picking among the K, so you only get bs if you want bs. I occasionally get people telling me AI is unreliable, and I tell them the same thing: the tech is nearly infinitely flexible (computing over the space of ideas!), so that says a lot more about how they’re using it.

So to really think the person writing the paper is the right one to review it?

Re: Major AI conference flooded with peer reviews written by AI

#136

Earlier quoted context omitted.

It's because people are commodities now. Human resources exists to manage the shuffle between warm bodies. It's back to OP's point. There's no such thing as professions now. Just jobs. We put them on and off like hats. With that churn comes lack of institutional knowledge and a rule set handed down from the C Suite for front line employees completely detached from the front line work. Enshitification run rampant.

But even given that, how is it that everything doesn't work very well? The normal functioning of markets would be that badly-working things are slowly driven out, while well-working things grow and replace them. Even without any reference to financial markets, this is simply what you expect to happen when people have a variety of things to choose from. I could hypothesize that markets have evolved to the point where…

Everything has always been "shit"...

Things might be shit in interesting and scary new ways, but there is no such thing as "the good ole days". Our mind wants to believe that things could go back to "how they used to be", "when it was better" but it's a fantasy.

It's an inability to face the cognitive dissonance and accept things as they are -- which is different than what we wanted! Boo hoo.

We all do this constantly everyday, some more than others :)

That said, humans are quite good at getting by even when things are shit. We've been doing it for untold eons.

Perhaps the only thing more impressive is how good we are at complaining about it all! Heh.

Re: Major AI conference flooded with peer reviews written by AI

#137
post #53

Earlier quoted context omitted.

Nothing points out that the benchmark is invalid like a zero false positive rate. Seemingly it is pre-2020 text vs a few models rework of texts. I can see this model fall apart in many real world scenarios. Yes, LLMs use strange language if left to their own devices and this can surely be detected. 0% false positive rate under all circumstances? Implausible.

Our benchmarks of public datasets put our FPR roughly around 1 in 10,000. https://www.pangram.com/blog/all-about-false-positives-in-ai... Find me a clean public dataset with no AI involvement and I will be happy to report Pangram's false positive rate on it.

Max, there's two problems I see with your comment.

1) the paper didn't show a 0% FNR. I mean tables 4, 7, and B.2 are pretty explicit. It's not hard to figure out from the others either.

2) a 0% error rate requires some pretty serious assumptions to be true. For that type of result to not be incredibly suspect requires there to be zero noise in the data, analysis, and at all parts. I do not see that being true of the mentioned dataset.

Even high scores are suspect. Generalizing the previous a score is suspect if it is higher than the noise level. Can you truly attest that this condition is true?

I'm suspect that you're introducing data leakage. I haven't looked enough into your training and data to determine how that's happening but you'll probably need a pretty deep analysis as leakage is really easy to sneak in. It can do so in non obvious ways. A very common one is tuning hyper parameters on test results. You don't have to pass data to pass information. Another sly way for this to happen is that the test set isn't significantly disjoint from the training set. If the perturbation is too small then you aren't testing generalization you're testing a slightly noisy training set (which your training should be introducing noise to help regularize, so you end up just measuring your training performance).

Your numbers are too good and that's suspect. You need a lot more evidence to suggest they mean what you want them to mean.

Re: Major AI conference flooded with peer reviews written by AI

#138

Whether it’s actually 20% or not doesn’t matter, everyone is aware the signal of the top confs is in freefall. There are also rings of reviewer fraud going on where groups of people in these niche areas all get assigned their own papers and recommend acceptance and in many cases the AC is part of this as well. Am not saying this is common but it is occurring. It feels as if every layer of society is in maximum extrac…

> spending time to carefully and deeply review a paper because they care and they feel on principal that’s the right thing to do

Generally agree, although several parts of that issue.

One of the first was covered by a paper back in 2023 that speaks to the issue about maximum extraction mode. [1] Fairness, honesty, and loyalty are usually rewarded with exploitation. If you spend time to carefully and deeply review the paper, then that ironically marks you as someone that can be exploited. You're implicitly marked as someone who will make personal sacrifices for the academic community and allow even more awful behavior to be piled on top of you. Unless they're caught with something especially egregious, the people that don't, get promoted, spend less time on reviews, and get further rewards.

[1] https://www.sciencedirect.com/science/article/abs/pii/S00221...

The academic community has talked about this a bunch for years. Editors / reviewers that don't paid, or get minimal payment, and sacrifice large amounts of their personal time effectively volunteering, while authors pay $1000's for each paper submitted, and then journals charge $10,000's for each subscription. It's been talked about for decades, and yet in all that time, very little has actually occurred to change the situation.

Another part on top of the "deeply reviewing papers" is that the sheer volume has massively increased (which has been an issue in a bunch of industries, sci-fi compilation Clarkesworld broke for quite a while in 2023 for similar reasons [2]). In the land of "type a sentence, and get a free academic paper" the extremely prolific are pouring out a paper a month, sometimes greater amounts. In areas like clinical medicine, hyper-prolific publishing has hit 70+ papers a year rates. [3] ~1.5 papers a week. Every few days somebody cranks out yet another paper that needs to be reviewed. In the article linked, one author had 140 articles to a single journal alone. Almost 3 times a week, all year long, you've got a paper claiming research worthy of publishing you need to review.

[2] https://neil-clarke.com/how-ai-submissions-have-changed-our-...

[3] https://www.sciencedirect.com/science/article/pii/S175115772...

One that I have less direct, citeable proof for, yet am rather suspicious of, is that theft has also dramatically increased with a huge surge in invasive monitoring and snooping. If my TV changes what I'm watching, and what's recommended, because I typed a text message to somebody, it seems likely that a lot of academia is also dealing with massive intellectual theft issues. This then heavily prioritizes pouring out material as quickly as possible, with as little effort as possible, to get the equivalent of first post and maximal posts, before it can be scraped, exfiltrated, and published by somebody else.

Finally, a lot of the reward and incentive has become metric chasing. Publish or Perish [4] and the Replication Crisis [5] are relatively well known ideas. Citation is a proxy of the impact of a paper, tenure and advancement is heavily related to quantity of publications and citations, and researchers would prefer to be cited more. And weirdly, if it does not work, and it's junk work, in a theme with the above, then it has been suggested nonreplicable publications are cited more than replicable ones [6]. In the linked paper, the view is that when "interesting" findings are published, they get more views, more media, more citations, and lower review standards get applied. And afterward there's very little social punishment for proving the results are false and not replicable (or reward for those illustrating lack of reproducability). Notably, the paper actually got a counterpoint stating that in psychology at least, lack of replication eventually predicts citation decline [7] (cited by 10), while the original actually got its authors ~250 citations, and a bunch of media mentions.

[4] https://en.wikipedia.org/wiki/Publish_or_perish

[5] https://en.wikipedia.org/wiki/Replication_crisis

[6] https://www.science.org/doi/10.1126/sciadv.abd1705

[7] https://www.pnas.org/doi/10.1073/pnas.2304862120

Re: Major AI conference flooded with peer reviews written by AI

#139

Earlier quoted context omitted.

But even given that, how is it that everything doesn't work very well? The normal functioning of markets would be that badly-working things are slowly driven out, while well-working things grow and replace them. Even without any reference to financial markets, this is simply what you expect to happen when people have a variety of things to choose from. I could hypothesize that markets have evolved to the point where…

Everything has always been "shit"... Things might be shit in interesting and scary new ways, but there is no such thing as "the good ole days". Our mind wants to believe that things could go back to "how they used to be", "when it was better" but it's a fantasy. It's an inability to face the cognitive dissonance and accept things as they are -- which is different than what we wanted! Boo hoo. We all do this constantl…

Your comment is a thought-terminating comment.

Re: Major AI conference flooded with peer reviews written by AI

#140

Earlier quoted context omitted.

Everything has always been "shit"... Things might be shit in interesting and scary new ways, but there is no such thing as "the good ole days". Our mind wants to believe that things could go back to "how they used to be", "when it was better" but it's a fantasy. It's an inability to face the cognitive dissonance and accept things as they are -- which is different than what we wanted! Boo hoo. We all do this constantl…

Your comment is a thought-terminating comment.

[deleted]
Post reply on HN