How we measured AI writing across arXiv, and where the measurement breaks
21–30 of 185 posts
Re: How we measured AI writing across arXiv, and where the measurement breaks
#22Re: How we measured AI writing across arXiv, and where the measurement breaks
#23The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…
My problem therefore is: we are seeing more and more papers written with tools that are known to make up facts, citations, and even entire papers. And the number of papers has increased, too. I therefore see it less as "people are being more productive" and more "people are releasing bad science much faster than we can keep up with".
Re: How we measured AI writing across arXiv, and where the measurement breaks
#24The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…
If fake papers weren't already a big problem before AI and the fields had already been policing themselves adequately, if this was already a functioning high-trust domain, maybe we could ignore this a bit more, but the fields already manifestly had problems. People taking advantage of that are reasonably more likely to use AI. The pressures to publish or perish provide the voltage and the AIs are a rather convenient path-to-ground.
I agree in some sense that if a truth is published, it doesn't matter if the AI or a human published it. However there are perfectly reasonable reasons to be concerned that AI usage is correlated to not publishing truths, especially in a world where merely being human-generated was already not an adequate check against that.
Re: How we measured AI writing across arXiv, and where the measurement breaks
#25Awesome work, the detector seems accurate on a bunch of texts I tried, surprisingly even human/ai hybrid texts
Re: How we measured AI writing across arXiv, and where the measurement breaks
#26> If a tool marks 40% of new papers as machine-written but also marks 20% of papers written before ChatGPT existed, the real story is the 20% nobody mentioned. When 65% of the papers you read have the characteristics of being AI written, whether or not you use AI to write, your writing will be influenced by the AI style. I imagine this must be particularly the case for newbie researchers who are still developing thei…
To me it sounds like 1. Either your tool is just not that good and reliable as you thought, 2. AI is trained on human written articles, so some of that human written content informed the now established “AI slop”.
There are people who shipped “slop” before AI.
Re: How we measured AI writing across arXiv, and where the measurement breaks
#27I scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally…
Re: How we measured AI writing across arXiv, and where the measurement breaks
#28The question is just how to organize these outputs and conclusions in a way that is consistently reproducible and also how to correct errors or remove LLM nonsense where it refuses to take a position on something.
Before it made sense to do this in papers but it feels like we need something like a paper format.. that is fully reproducible ideally and optimized for aggregating knowledge in a better way. I.e. before a person spent months on one of these and there was just more filtering, and the output itself was a clear signal of time spent and effort that no longer exists.
Re: How we measured AI writing across arXiv, and where the measurement breaks
#29The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…
Re: How we measured AI writing across arXiv, and where the measurement breaks
#30Weimar moment of science