Live data from Hacker News

Major AI conference flooded with peer reviews written by AI

nature.com

111–120 of 140 posts

Re: Major AI conference flooded with peer reviews written by AI

#111

Earlier quoted context omitted.

There are dozens of first generation AI detectors and they all suck. I'm not going to defend them. Most of them use perplexity based methods, which is a decent separators of AI and human text (80-90%) but has flaws that can't be overcome and high FPRs on ESL text. https://www.pangram.com/blog/why-perplexity-and-burstiness-f... Pangram is fundamentally different technology, it's a large deep learning based model that…

Can your software detect which LLMs most likely generated a text?

Pangram is trained on this task as well to add additional signal during training, but it's only ~90% accurate so we don't show the prediction in public-facing results

Re: Major AI conference flooded with peer reviews written by AI

#112
post #78

Earlier quoted context omitted.

Don't kid yourself, all those steps have AI heavily involved in them. And that's not necessarily a bad thing. If I set up RAG correctly, then tell the AI to generate K samples, then spend time to pick out the best one, that's still significant human input, and likely very good output too. It's just invisible what the human did. And as models get better, the necessary K will become smaller....

That's a strategy for producing maximally convincing BS content, but the scientific method was absent. I'm sure it was an oversight... ; )

That’s on you. You get to decide what “best” means when picking among the K, so you only get bs if you want bs.

I occasionally get people telling me AI is unreliable, and I tell them the same thing: the tech is nearly infinitely flexible (computing over the space of ideas!), so that says a lot more about how they’re using it.

Re: Major AI conference flooded with peer reviews written by AI

#113

Earlier quoted context omitted.

I am not sure if you are familiar with Pangram (co-founder here) but we are a group of research scientists who have made significant progress in this problem space. If your mental model of AI detectors is still GPTZero or the ones that say the declaration of independence is AI, then you probably haven't seen how much better they've gotten. This paper by economists from the University of Chicago economists found zero…

Are you concerned with your product being used to improve AI to be less detectable?

It's definitely going to be a back and forth - model providers like OpenAI want their LLMs to sound human-like. But this is the battle we signed up for, and we think we're more nimble and can iterate faster to stay one step ahead of the model providers.

Re: Major AI conference flooded with peer reviews written by AI

#114
post #80

Earlier quoted context omitted.

Yeah, Pangram does not provide any concrete proof, but it confirms many people's suspicions about their reviews. But it does flag reviews for a human to take a closer look and see if the review is flawed, low-effort, or contains major hallucinations.

Was there an analysis of flawed, low-effort reviews in similar conferences before generative AI models? From what I remember, (long before generative AI) you would still occasionally get very crappy reviews (as author). When I participated (couple of times) to review committees, when there was a high variance between reviews the crappy reviews were rather easy to spot and eliminate. Now it's not bad to detect crappy…

Anecdotally people are seeing a rise of low-quality reviews which is correlated with increased reviewer workload and and AI tools giving reviews an easy way out. I don't know of any studies quantifying review quality, but I would recommend checking the Peer Review Congress program from past years.

Re: Major AI conference flooded with peer reviews written by AI

#115

Earlier quoted context omitted.

Are you concerned with your product being used to improve AI to be less detectable?

It's definitely going to be a back and forth - model providers like OpenAI want their LLMs to sound human-like. But this is the battle we signed up for, and we think we're more nimble and can iterate faster to stay one step ahead of the model providers.

That sounds extremely naive but good luck!

Re: Major AI conference flooded with peer reviews written by AI

#116
post #53

Earlier quoted context omitted.

Nothing points out that the benchmark is invalid like a zero false positive rate. Seemingly it is pre-2020 text vs a few models rework of texts. I can see this model fall apart in many real world scenarios. Yes, LLMs use strange language if left to their own devices and this can surely be detected. 0% false positive rate under all circumstances? Implausible.

Our benchmarks of public datasets put our FPR roughly around 1 in 10,000. https://www.pangram.com/blog/all-about-false-positives-in-ai... Find me a clean public dataset with no AI involvement and I will be happy to report Pangram's false positive rate on it.

[flagged]

Re: Major AI conference flooded with peer reviews written by AI

#117
post #91

Earlier quoted context omitted.

Just yesterday I saw this YouTube rant from someone called Jaiden Animations, about how everything is just shit now. https://www.youtube.com/watch?v=NBZv0_MImIY She opens with an example of a bank. She walked in and asked for a debit card. The teller told her to take a seat. 30 minutes later, the teller told her the bank doesn't issue debit cards. Firstly, what kind of bank doesn't issue debit cards, and secondly, wh…

It's because people are commodities now. Human resources exists to manage the shuffle between warm bodies. It's back to OP's point. There's no such thing as professions now. Just jobs. We put them on and off like hats. With that churn comes lack of institutional knowledge and a rule set handed down from the C Suite for front line employees completely detached from the front line work. Enshitification run rampant.

But even given that, how is it that everything doesn't work very well?

The normal functioning of markets would be that badly-working things are slowly driven out, while well-working things grow and replace them. Even without any reference to financial markets, this is simply what you expect to happen when people have a variety of things to choose from.

I could hypothesize that markets have evolved to the point where it's impossible for new things to grow unless they are already shit. Perhaps because everyone's too busy working for the shit things (which is partly because the government keeps printing money to the previously successful things in order to prevent the economy collapsing and therefore landlords got to charge exorbitant rent) or perhaps because they just don't have any money because of the above, and can only afford the cheap shit things (but a lot of the shit things are expensive?) or perhaps because people are afraid to start new things because they're afraid of the government (I've observed that not infrequently on HN, also something something testosterone microplastics) or perhaps because advertising effectiveness has reached the point where new things never become discoverable and stay crowded out as old things ramp up advertisement to compensate or perhaps we're just all depressed (because of the housing market probably).

Re: Major AI conference flooded with peer reviews written by AI

#119
Sorry to say but it's another example of the destructive power of AI, along the lines of no longer being able to establish "truth" now that any evidence (video, audio, image, etc.) can be explicitly faked (yes, AI detectors exist but that will be a continuous race with AIs designed to outsmart the detectors). The end result could be that peer reviews become worthless and trust in scientific research -- already at an all time low -- becomes even lower. Sad.

Re: Major AI conference flooded with peer reviews written by AI

#120
post #6

While I think there's significant AI "offloading" in writing, the article's methodology relies on "AI-detectors," which reads like PR for Pangram. I don't need to explain why AI detectors are mostly bullshit and harmful for people who have never used LLMs. [1] 1: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

I am not sure if you are familiar with Pangram (co-founder here) but we are a group of research scientists who have made significant progress in this problem space. If your mental model of AI detectors is still GPTZero or the ones that say the declaration of independence is AI, then you probably haven't seen how much better they've gotten. This paper by economists from the University of Chicago economists found zero…

How do you discern between papers "completely fabricated" by AI vs. edited by AI for grammar?
Post reply on HN