I scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally…
How we measured AI writing across arXiv, and where the measurement breaks
111–120 of 185 posts
Re: How we measured AI writing across arXiv, and where the measurement breaks
#112I don't think the problem is as bad as a naive reading of this article suggests. I'm highly skeptical that anywhere near 65% of recent CS papers that I've read (mostly systems papers) are substantially AI-written. I threw some recent papers I've read into the system and they come back as 0-7%.
Re: How we measured AI writing across arXiv, and where the measurement breaks
#113People are saying this is a bad thing but is it really a problem? The compelling aspect of research is the data and/or description of work, not the writing. Papers probably should be written by AI so that they're clear and well presented, while the researchers should focus on generating good data. If there is no data or work behind the paper, we should question whether the research group needs funding.
at the least, this is problematic for peer-review because the absolute number of submissions outpaces the time availability of a finite number of expert reviewers. We cannot quickly generate expert human reviewers, and so the community might converge towards half-baked solutions (AI-generated reviews or rejection systems, vastly expanded referee pools, etc.) that tend to erode trust and and make scientific communities more adversarial.
Re: How we measured AI writing across arXiv, and where the measurement breaks
#114The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…
LLM-written text tends to have the property that it gives the impression of expertise and knowledge in excess of what the text actually contains. In other words, it "sounds smart" without necessarily having anything to back it up. In even more critical terms, it's very good at bullshitting. Unfortunately for us, the scientific community current relies on a certain amount of trust. (To do otherwise is very expensive!…
Re: How we measured AI writing across arXiv, and where the measurement breaks
#115Just to play devils advocate. These kind of papers are very verbose and boilerplate. I can imagine using AI to write 90% but then the actual novel content and explaining what’s important could be handwritten. Perhaps that’s what’s happening.
If 90% of it is boilerplate, then you really need to question if you have something worth publishing as an academic paper.
Re: How we measured AI writing across arXiv, and where the measurement breaks
#116I'm not sure what to think... I uploaded a PyHPC workshop paper I wrote in 2011 and it said 27% machine. I also uploaded my PhD dissertation from 2012 and got back 40% machine, which is just barely below the 42% threshold. I don't publish anymore... but does this mean I wrote like a LLM or did LLMs learn from me? :p Update: I also uploaded a IEEE CLUSTERS paper I wrote in 2015 and it came back 74% machine written :|
Re: How we measured AI writing across arXiv, and where the measurement breaks
#117Re: How we measured AI writing across arXiv, and where the measurement breaks
#118Honestly academic writing is the only place where I think AI slop might be an improvement over the status quo writing style.
Even if they manage to avoid straight-up factual incorrectness, the writing is full of ambiguities and vagueness when you look closely. Meanwhile, space is wasted repeating the same claim in multiple ways, or explaining something simple.
They also seem unable to resist the hype/advertising tone, overselling the contribution while exaggerating the limitations of related work.
I would much rather read grammatically incorrect or awkward sentences.
While academic writing does have a few pointless historical conventions, the huge majority of "status quo writing style" is a logical consequence of 1) minimizing ambiguity, 2) organizing ideas coherently, 3) distinguishing opinion/interpretation from fact, and 4) providing enough detail to be reproducible.
Re: How we measured AI writing across arXiv, and where the measurement breaks
#119The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…
Another important questions is: What about the papers that graduate to proper publication? Arxiv is full of pre-prints that anyone can upload.
You now (at least for some categories) have to receive endorsement from someone who has multiple recent papers on arxiv in the same (or adjacent) category.