Live data from Hacker News

How we measured AI writing across arXiv, and where the measurement breaks

unslop.run

61–70 of 185 posts

Re: How we measured AI writing across arXiv, and where the measurement breaks

#61

I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each pa…

My English is fine, and I did hand-write my MSc thesis in English, but scientific style writing is very tedious. Personally I don't think I'd even consider writing a paper "by hand" if I were to write another one.

There really is no point, as long as you verify the content matches your intent and edit out anything poorly written.

Frankly, I've read plenty of papers by native English-speakers over the years that'd strongly benefit from being rewritten by an LLM too...

Re: How we measured AI writing across arXiv, and where the measurement breaks

#62
post #57
post #56

Earlier quoted context omitted.

> If you write each paragraph and have an LLM make that paragraph more scientific, it is entirely your paper "More scientific" is not some merely stylistic thing that faithfully preserves the original meaning of what you wrote. The precise details of each paragraph matters a lot in terms of what and how it communicates. The fact that these details do matter means that, according to my accounting, it is not entirely y…

You might prefer that, but journal editors might not.

I've never had any trouble publishing in a more conversational style and have read plenty of papers that do the same. Of course, this varies by field and venue and format (journal vs. conference) and even reviewer and it's difficult to understand what the true constraints are and what the state of the publishing system is in this regard.

But I suspect that a lot of academic's feelings about it are informed by what others have told them and how they've been trained, rather than by what's actually permissible in the publishing system.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#63
post #59

Earlier quoted context omitted.

[flagged]

Which part is a sin? Using LLMs to deal with a lack of English-language fluency? I am a scientist (actually a mathematician, if it matters), and, if that's the way to deal with the practical hegemony of English in the scientific literature, then I have no problem with it. Rather that than people with important ideas can't get them before the scientific community. As long as the authors personally check and stand behi…

[flagged]

Re: How we measured AI writing across arXiv, and where the measurement breaks

#64
post #16

Earlier quoted context omitted.

Even if you can reject the aesthetic argument for non-fiction works (although read some of Dijkstra's papers for a good counterargument), it is still a problem because it breaks an important quality signalling mechanism. Pre-LLMs, a paper with no spelling or grammar errors showed that somebody had put effort into writing and editing it. If they cared about the presentation, they probably also cared about the content.…

This is fallacious. The fact that someone used LLM to write their paper does not negate its findings.

That's not the point though, the point being made ist that an LLM written paper very likely doesn't actually find something despite looking like it does on the surface

Re: How we measured AI writing across arXiv, and where the measurement breaks

#65
I'm not sure what to think... I uploaded a PyHPC workshop paper I wrote in 2011 and it said 27% machine.

I also uploaded my PhD dissertation from 2012 and got back 40% machine, which is just barely below the 42% threshold.

I don't publish anymore... but does this mean I wrote like a LLM or did LLMs learn from me? :p

Update: I also uploaded a IEEE CLUSTERS paper I wrote in 2015 and it came back 74% machine written :|

Re: How we measured AI writing across arXiv, and where the measurement breaks

#66

> If a tool marks 40% of new papers as machine-written but also marks 20% of papers written before ChatGPT existed, the real story is the 20% nobody mentioned. When 65% of the papers you read have the characteristics of being AI written, whether or not you use AI to write, your writing will be influenced by the AI style. I imagine this must be particularly the case for newbie researchers who are still developing thei…

I might be missing something but what’s the real story of the 20%. To me it sounds like 1. Either your tool is just not that good and reliable as you thought, 2. AI is trained on human written articles, so some of that human written content informed the now established “AI slop”. There are people who shipped “slop” before AI.

> There are people who shipped “slop” before AI.

The funny thing is that "slop" was defined by the writing habits of AI model, which we have learned to pick upon and recognize.

The "It's not X, it's Y", the rhetorical questions and other patterns would have been the tools of a skilled writer, and those people writing "like AI" before AI most likely would have been recognized as such.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#67

I scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally…

What detector are you using? How can you be sure of its accuracy given that every commercial AI detector has been debunked?

Re: How we measured AI writing across arXiv, and where the measurement breaks

#68

I scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally…

[deleted]

Re: How we measured AI writing across arXiv, and where the measurement breaks

#69
post #53

Just to play devils advocate. These kind of papers are very verbose and boilerplate. I can imagine using AI to write 90% but then the actual novel content and explaining what’s important could be handwritten. Perhaps that’s what’s happening.

Yes, this is the thing. In the scientific enterprise, the writing is mostly wasted time. Why would you write it up if machines can do it competently?

Sure, some people are artists - but most aren't.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#70
One challenge with this approach is: could it be possible that the detector is simply learning to recognize words and jargon used more in the literature post 2022 as 'AI'?

For example, LLMs love to talk about LLMs (and the people who write with LLMs love to write about LLMs). Could "large language model" itself therefore be flagged as an AI-like phrase by this approach? It didn't exist much in the literature before 2022, does now, and certainly does more in AI-generated text: but, it is not actually a great way to distinguish modern AI generated text from human written text.

A helpful control would be to show that on some cohort of papers that can be declared reasonably clean of LLM generated text post 2023 there are very low rates compared to the arxiv.

For example, while papers in the journals Nature and Science are unlikely to be entirely LLM free at this point, if those were tested through 2026, we should see a line significantly lower than the arxiv's growth.

Post reply on HN