Live data from Hacker News

How we measured AI writing across arXiv, and where the measurement breaks

unslop.run

151–160 of 185 posts

Re: How we measured AI writing across arXiv, and where the measurement breaks

#151
post #65

I'm not sure what to think... I uploaded a PyHPC workshop paper I wrote in 2011 and it said 27% machine. I also uploaded my PhD dissertation from 2012 and got back 40% machine, which is just barely below the 42% threshold. I don't publish anymore... but does this mean I wrote like a LLM or did LLMs learn from me? :p Update: I also uploaded a IEEE CLUSTERS paper I wrote in 2015 and it came back 74% machine written :|

When "detectors" first started popping up all over the place, all of them rated the declaration of independence as 100% AI written, so... Yeah, these things just don't work. And what's even more dangerous is that people that don't understand how any of it works use these tools, and accuse people of using AI, sometimes with grave consequences. Students have been through this, at all levels of education.

In fairness, the original crop of detectors were based on perplexity scores and were entirely useless (famously, the declaration of independence often came back as 100% AI generated).

I'm not convinced we'll ever have full-proof detectors, and certainly the false-positive rate will make them irresponsible for accusations of intellectual/academic fraud, I do think that LLMs are easy for folks to sniff out on average so I imagine it's possible to detect many instances.

Pangram's detector is anecdotally very accurate in my tests. This detector appears to be fine-tuned on a very small dataset (200 papers per subject), and suspect the problem might be in part that.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#152

I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each pa…

It seems that LLMs often insert new ideas even when asked only to polish writing without changing it: https://bsky.app/profile/tomerullman.bsky.social/post/3mq33c...

I don't think this is surprising. Good technical writing is very precise. If you're starting from non-technical writing, I suspect that in most cases you can't make it "sound scientific" without adding new claims or changing the meaning. (Maybe you are being more careful, but this is something that worries me in general.)

Re: How we measured AI writing across arXiv, and where the measurement breaks

#153

I scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally…

A signal, I put an AI generated of mine (with heavy human guidance on aesthetics mostly and some minor human editing) and got 5% only.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#154
post #67

Earlier quoted context omitted.

What detector are you using? How can you be sure of its accuracy given that every commercial AI detector has been debunked?

Almost every. Pangram is pretty accurate on longer texts. I haven't seen any glaring examples of false positives.

I’ve heard good things about Pangram but it’s trivially easy to fool it. I generated text using Opus 4.8 and then humanized it using Grammarly. Pangram determined it 100% human-written. Tried it numerous times with the same result. GPTZero claimed with 80% confidence that it was human-written but AI polished. Nothing in that text was human-written.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#155

Earlier quoted context omitted.

When "detectors" first started popping up all over the place, all of them rated the declaration of independence as 100% AI written, so... Yeah, these things just don't work. And what's even more dangerous is that people that don't understand how any of it works use these tools, and accuse people of using AI, sometimes with grave consequences. Students have been through this, at all levels of education.

I replied to a comment with a structured informal proof on HN. A well established user here was adamant that I used AI because apparently humans never ever wrote proofs. This was a while ago. Any well crafted human output is now being dismissively cast as AI if the reader is challenged by the output intellectually/politically.

It's the irl version of being accused of hacking in Counterstrike.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#156
post #59

Earlier quoted context omitted.

[flagged]

Which part is a sin? Using LLMs to deal with a lack of English-language fluency? I am a scientist (actually a mathematician, if it matters), and, if that's the way to deal with the practical hegemony of English in the scientific literature, then I have no problem with it. Rather that than people with important ideas can't get them before the scientific community. As long as the authors personally check and stand behi…

https://lists.isocpp.org/std-proposals/att-0486/Reply_to_Zer...

Here's a paper by a non native English speaker. The lack of the formal style doesn't cause any issue. Preciseness is what matters.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#157

I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each pa…

most papers like this are flagged by AI detectors because, by your own admission, AI was used to generate the final output. what exactly did you expect here?

Re: How we measured AI writing across arXiv, and where the measurement breaks

#158
post #141

Earlier quoted context omitted.

They use a "threshold calibrated so pre-ChatGPT papers flag at 0.4%", so these things do work most of the time. It also means that there are known false positives, so for any given paper, scoring above the threshold isn't irrefutable proof of AI usage. But for things like estimating the overall proportion of AI writing, you only need to be correct on average, so individual false positives don't matter much.

It would be interesting to see how that .4% varies across fields on arXiv. I imagine some areas influenced LLM writing much more than others and are more susceptible to false positives. Anyway it's a little strange to me that these detectors do so poorly (at least by what I hear on this site). There is as much labeled data as you would need to train on pre-LLM human vs LLM text. If their out of sample errors are as g…

There's a table in the article breaking down the rate by field.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#159

I scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally…

Couple of methodological notes

* pre-chatGPT is not an effective control because language evolves. In the arxiv corpus in particular there are "fashions" in research depending on what gets funded lately, not to mention many new words and topics not invented before a given paper.

* In general, detecting AI from content seems difficult as humans write like they read. To the extent there are unique factors recognizable as AI and to the extent humans read them, they will eventually incorporate them into their writing style. Accordingly, you'd need to model a rolling window of "AI tells" that decay at some rate.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#160
post #65

I'm not sure what to think... I uploaded a PyHPC workshop paper I wrote in 2011 and it said 27% machine. I also uploaded my PhD dissertation from 2012 and got back 40% machine, which is just barely below the 42% threshold. I don't publish anymore... but does this mean I wrote like a LLM or did LLMs learn from me? :p Update: I also uploaded a IEEE CLUSTERS paper I wrote in 2015 and it came back 74% machine written :|

It means machine text and human text cannot be distinguished from each other.

It's just text.

Post reply on HN