Live data from Hacker News

How we measured AI writing across arXiv, and where the measurement breaks

unslop.run

11–20 of 185 posts

Re: How we measured AI writing across arXiv, and where the measurement breaks

#11
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

The problem is that humans start with credulity. AI hallucinates and makes up stuff some percentage of the time. Humans are not "default deny" when given information. Unleashing that was a disservice to mankind and has created a future filled with lies and people who are confident in them.

Oh, to be technically correct:

AI hallucinates and makes up stuff 100% percent of the time. Never been a fan of that word for this.

Again, I fail to see the problem here that isn't solved by careful reading WHICH IS WHAT PEOPLE SHOULD BE DOING ANYWAY. I would like to see room for AI disclosure, maybe a statement of "this is how much AI I used."

But this blanket X% of this looks like AI? Again, so what?

Re: How we measured AI writing across arXiv, and where the measurement breaks

#12

Earlier quoted context omitted.

Surprisingly I agree with you. My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

I mean, sure. But what an absolutely insane predicate. "Not reducing the quality of the output substantially" is (as an AI might say...) load bearing there. Other problems include: Signal to Noise Ratio going through the roof.

Yeah I was careful on purpose with my statement. :D

Re: How we measured AI writing across arXiv, and where the measurement breaks

#13

Earlier quoted context omitted.

Surprisingly I agree with you. My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

> My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. That's a big if. ArXiv is not peer reviewed and LLMs basically interpolate and extrapolate text, which makes them essentially fluff generators. Even in the most charitable interpretation, LLMs enable those with nothing to say to say nothing while meeting surface-level style g…

We can't know if real science is happening in the background but I'd wager that the majority of these papers is not complete slop but real findings with AI generated text used to communicate it. If it was just straight slop I would be really worried.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#14
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

you just dont respect your readers, thats all

Re: How we measured AI writing across arXiv, and where the measurement breaks

#15

Earlier quoted context omitted.

Surprisingly I agree with you. My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

I mean, sure. But what an absolutely insane predicate. "Not reducing the quality of the output substantially" is (as an AI might say...) load bearing there. Other problems include: Signal to Noise Ratio going through the roof.

Is it? I guess look -- I'm an academic, I've read piles of articles such as these. Signal to noise, already not great.

Yes, I feel like there's room to improve things, I just strongly doubt that "using AI to detect AI" is a particularly useful thing to do here.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#16
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

Even if you can reject the aesthetic argument for non-fiction works (although read some of Dijkstra's papers for a good counterargument), it is still a problem because it breaks an important quality signalling mechanism.

Pre-LLMs, a paper with no spelling or grammar errors showed that somebody had put effort into writing and editing it. If they cared about the presentation, they probably also cared about the content. LLMs routinely produce nonsense that looks superficially like high-quality work.

There are far too many papers to read all of them. LLM slop is evidence that something is probably low quality. As the saying goes, "if you can't be bothered writing it, I can't be bothered reading." The rare outliers will get enough citations and recommendations to overcome this filter.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#17
> If a tool marks 40% of new papers as machine-written but also marks 20% of papers written before ChatGPT existed, the real story is the 20% nobody mentioned.

When 65% of the papers you read have the characteristics of being AI written, whether or not you use AI to write, your writing will be influenced by the AI style. I imagine this must be particularly the case for newbie researchers who are still developing their writing style

Re: How we measured AI writing across arXiv, and where the measurement breaks

#18
post #16
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

Even if you can reject the aesthetic argument for non-fiction works (although read some of Dijkstra's papers for a good counterargument), it is still a problem because it breaks an important quality signalling mechanism. Pre-LLMs, a paper with no spelling or grammar errors showed that somebody had put effort into writing and editing it. If they cared about the presentation, they probably also cared about the content.…

This is fallacious. The fact that someone used LLM to write their paper does not negate its findings.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#19

Earlier quoted context omitted.

> My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. That's a big if. ArXiv is not peer reviewed and LLMs basically interpolate and extrapolate text, which makes them essentially fluff generators. Even in the most charitable interpretation, LLMs enable those with nothing to say to say nothing while meeting surface-level style g…

We can't know if real science is happening in the background but I'd wager that the majority of these papers is not complete slop but real findings with AI generated text used to communicate it. If it was just straight slop I would be really worried.

> (...) but I'd wager that the majority of these papers is not complete slop (...)

That's a huge assumption, and one that goes against the whole notion of using LLMs to generate text. AI slop is by far the norm.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#20
post #16

Earlier quoted context omitted.

Even if you can reject the aesthetic argument for non-fiction works (although read some of Dijkstra's papers for a good counterargument), it is still a problem because it breaks an important quality signalling mechanism. Pre-LLMs, a paper with no spelling or grammar errors showed that somebody had put effort into writing and editing it. If they cared about the presentation, they probably also cared about the content.…

This is fallacious. The fact that someone used LLM to write their paper does not negate its findings.

That's not the point being made.
Post reply on HN