Live data from Hacker News

How we measured AI writing across arXiv, and where the measurement breaks

unslop.run

71–80 of 185 posts

Re: How we measured AI writing across arXiv, and where the measurement breaks

#71
post #65

I'm not sure what to think... I uploaded a PyHPC workshop paper I wrote in 2011 and it said 27% machine. I also uploaded my PhD dissertation from 2012 and got back 40% machine, which is just barely below the 42% threshold. I don't publish anymore... but does this mean I wrote like a LLM or did LLMs learn from me? :p Update: I also uploaded a IEEE CLUSTERS paper I wrote in 2015 and it came back 74% machine written :|

It's been trained on a work and then distilled, the false positive noise has to be absurd.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#72
post #40

As with all text-only AI detection schemes, I am concerned about the accuracy of the detection. I'm skeptical of the methodology, specifically the final join of the three detector scores---how can you be sure that that final step does not introduce any biases? There's no source available, so it's difficult to tell exactly how this works or reproduce the research. I've also uploaded text samples from my own (unrelease…

The most concerning is the use of these tools to detect cheaters in schools.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#73
People are saying this is a bad thing but is it really a problem? The compelling aspect of research is the data and/or description of work, not the writing. Papers probably should be written by AI so that they're clear and well presented, while the researchers should focus on generating good data. If there is no data or work behind the paper, we should question whether the research group needs funding.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#75
post #67

I scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally…

What detector are you using? How can you be sure of its accuracy given that every commercial AI detector has been debunked?

Almost every. Pangram is pretty accurate on longer texts. I haven't seen any glaring examples of false positives.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#76

One challenge with this approach is: could it be possible that the detector is simply learning to recognize words and jargon used more in the literature post 2022 as 'AI'? For example, LLMs love to talk about LLMs (and the people who write with LLMs love to write about LLMs). Could "large language model" itself therefore be flagged as an AI-like phrase by this approach? It didn't exist much in the literature before 2…

You raise a valid point. Just by intuition I'd say if this were true, it would probably just be a small fraction of the actual flagged articles. I will still look into how I can mitigate this when I update the detector.

The difficulty with this is then: How do you get a clean post 2023 dataset? I have no straightforward idea for this. You can't use other AI detectors to build it because then you'd never outperform them.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#77
post #53

Just to play devils advocate. These kind of papers are very verbose and boilerplate. I can imagine using AI to write 90% but then the actual novel content and explaining what’s important could be handwritten. Perhaps that’s what’s happening.

If 90% of it is boilerplate, then you really need to question if you have something worth publishing as an academic paper.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#78
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

I recently desk-rejected a paper where every single citation in its Introduction was hallucinated. That means that the entire connection between what the author(s) did and how it relates to existing research was simply made up. I've never seen this happening before AI but now there's at least one paper in every cycle pulling something similar. My problem therefore is: we are seeing more and more papers written with t…

Yeah but that's an issue with the researcher putting out a bad paper, and it suggests you'll have to reject more papers. We wouldn't ban email because many of the emails are spam, it just means we need new tools to filter out junk. AI will allow researchers to be more productive all together and take less time to publish a paper, which is good.
Post reply on HN