Live data from Hacker News

How we measured AI writing across arXiv, and where the measurement breaks

unslop.run

91–100 of 185 posts

Re: How we measured AI writing across arXiv, and where the measurement breaks

#91
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

Surprisingly I agree with you. My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

> if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue.

That's a big "If".

If a research is good, the author still has to clear all the hurdles in publishing. "Writing your own paper" is just one more hurdle.

> A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

That's just a different way of saying "if the majority of CS papers are crap, lets just accept this reality".

So, go on, publish away all your AI-induced "research", but the bar is slowly going to be raised anyway to reject that. That's how science always worked - when a bar is not sufficient to exclude the crap, it is raised.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#92

Earlier quoted context omitted.

> My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. That's a big if. ArXiv is not peer reviewed and LLMs basically interpolate and extrapolate text, which makes them essentially fluff generators. Even in the most charitable interpretation, LLMs enable those with nothing to say to say nothing while meeting surface-level style g…

We can't know if real science is happening in the background but I'd wager that the majority of these papers is not complete slop but real findings with AI generated text used to communicate it. If it was just straight slop I would be really worried.

> We can't know if real science is happening in the background but I'd wager that the majority of these papers is not complete slop but real findings with AI generated text used to communicate it. If it was just straight slop I would be really worried.

Why don't you read them and see? The ones I looked at were clear slop.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#94
post #7

Earlier quoted context omitted.

This is a pretty stunning result. The time series looks really convincing. Is the way the detector itself is trained orthogonal to this or could there be some "leakage" in that the pre-chatgpt text is in the (positive) training data?

I tried my best to avoid leakage. If you're curious about how I trained the detector I have a writeup on it: https://unslop.run/blog/how-our-ai-text-detector-works FYI this is all relatively new so there might be lots of issues and iterations coming.

My honest first impressions, since I think the project is well intended: This writeup is itself AI, and I would venture to call it slop. The Calibration section is borderline uninterpretable, and I challenge any non-author who claims to understand it to answer some basic peer review questions about it.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#96
To corroborate this - I found pretty similar rates of AI-flagged papers over time from running pangram on ArXiv (abstracts) in a physics subfield, which is currently at about 25% as of April (write up here: https://peterse.github.io/2026/06/15/The-rising-tide-pt1.htm...).

Its not clear from your writeup what threshold needs to be reached to be classified as "machine written". A preprint where half the text is human and half is 100% AI should be a different category than a preprint where 100% of the text is AI-assisted.

Also its cool that you're making the detector available. When you say "cheap to run", do you know how this compares to pricing for a commercial detector pangram or GPTZero?

Re: How we measured AI writing across arXiv, and where the measurement breaks

#97
post #85

There some real game theory mechanics at play when it comes to LLM usage in corporations. Devs are cranking out superficially superior code and documentation by just aiming the Claude Code fire hose at everything they can. Leadership encourages this because from what they can tell, there is no downside. It's hard to say if this code is structurally better or worse than before, but it's certainly voluminous and as far…

> "superficially superior code"

Can you unpack that a bit? It produces measurably, meaningfully inferior code everywhere I see it in use.

> "leadership encourages this because from what they can tell, there is no downside"

As said 'leadership' I find this a bit puzzling. I'm seeing strong, quantifiable evidence of increasing churn, increasing incident count, and length of downtime from the date of our biggest push into GenAI, and I'm organizing efforts on my teams to mitigate those issues and actively reduce GenAI adoption.

If you mean my c-suite, you're mostly correct although they are already rumbling about seeing zero or negative ROI on GenAI investments.

> Anyone not using LLMs all day is just not going to be as prolific

Agreed, but prolific != productive.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#98

I run most professional statements and articles through LLMs before submission. This helps correct grammar and improves accessibility through better syntax (because I write exactly how I think). The article doesn't seem to mention consideration of AI for polishing human work.

AI doesn't polish human work. It is more like an extruder that squeezes anything you put into it into a generic, formulaic shape, indistinguishable from writing that was lazily prompted because the writer couldn't be bothered to put forth any more effort than that.

"AI does not polish human work. It acts more like an extruder, forcing anything fed into it into the same generic, formulaic shape—indistinguishable from writing produced by a lazy prompt from someone unwilling to put in any more effort."

There. AI-polished sentence.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#99

Earlier quoted context omitted.

Surprisingly I agree with you. My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

I think the key question is more effective at what? I see plenty of anecdotal evidence that models have been trained fantastically well—and getting better—at writing to trigger the right neurons in the human population to produce “This is interesting/informative/correct” responses in bulk. Could their ability to produce those responses run far ahead of their ability to actually achieve the last in reality? Sure seems…

Many people are bad at writing, including scientists. Improvements could be articulating a concept in a way that the reader will understand it clearly.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#100
post #16

Earlier quoted context omitted.

Even if you can reject the aesthetic argument for non-fiction works (although read some of Dijkstra's papers for a good counterargument), it is still a problem because it breaks an important quality signalling mechanism. Pre-LLMs, a paper with no spelling or grammar errors showed that somebody had put effort into writing and editing it. If they cared about the presentation, they probably also cared about the content.…

This is fallacious. The fact that someone used LLM to write their paper does not negate its findings.

It certainly casts suspicion on the ‘findings’ which may well be hallucinated or subtly wrong.
Post reply on HN