Live data from Hacker News

How we measured AI writing across arXiv, and where the measurement breaks

unslop.run

51–60 of 185 posts

Re: How we measured AI writing across arXiv, and where the measurement breaks

#51

I scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally…

this is neat! is the model available somewhere for local execution? or even a lookup table with your results for all arXiv pre-print codes? I want to run it on lots of pre-prints and I don't want to kill your server

Thank you, I plan to release the arxiv preprint codes. Also don't worry about killing my server, let me know if you succeed :D

Re: How we measured AI writing across arXiv, and where the measurement breaks

#52

I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each pa…

Yes that is exactly what I suspected. And I see absolutely nothing wrong with this.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#53
Just to play devils advocate. These kind of papers are very verbose and boilerplate. I can imagine using AI to write 90% but then the actual novel content and explaining what’s important could be handwritten.

Perhaps that’s what’s happening.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#54

the funniest part of these AI detectors is that if I were to upload any of einstin's paper's they will all be flagged as AI-written. it makes sense because it's part of their training data. but this post makes me wonder, if more papers' are written with AI, or the shape of knowledge of converging?

> the funniest part of these AI detectors is that if I were to upload any of einstin's paper's they will all be flagged as AI-written.

Have you tried doing that or even read the article?

The article says that their detector flags 0.4% of pre-AI papers as AI-written.

If I paste the first page from this paper (https://www.fourmilab.ch/etexts/einstein/specrel/specrel.pdf) in https://unslop.run/app, I get a 0% chance that it was AI-written.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#55
post #54

the funniest part of these AI detectors is that if I were to upload any of einstin's paper's they will all be flagged as AI-written. it makes sense because it's part of their training data. but this post makes me wonder, if more papers' are written with AI, or the shape of knowledge of converging?

> the funniest part of these AI detectors is that if I were to upload any of einstin's paper's they will all be flagged as AI-written. Have you tried doing that or even read the article? The article says that their detector flags 0.4% of pre-AI papers as AI-written. If I paste the first page from this paper ( https://www.fourmilab.ch/etexts/einstein/specrel/specrel.pdf ) in https://unslop.run/app , I get a 0% chance…

Even the translated version of the field equations paper has a 1% match only. Which I expected to be a tiny bit higher but still near 0.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#56

I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each pa…

> If you write each paragraph and have an LLM make that paragraph more scientific, it is entirely your paper

"More scientific" is not some merely stylistic thing that faithfully preserves the original meaning of what you wrote. The precise details of each paragraph matters a lot in terms of what and how it communicates. The fact that these details do matter means that, according to my accounting, it is not entirely your paper.

Also, I find striving for "scientific" to be a pretty undesirable thing. Why should papers read like that? What is the benefit? The best papers (in terms of their writing and communication) are unpretentious and conversational. I'm pretty sure I'd prefer your "basic sentences", especially if they were wholly yours. (I understand that there are also external forces at play here as you mentioned.)

Re: How we measured AI writing across arXiv, and where the measurement breaks

#57
post #56

I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each pa…

> If you write each paragraph and have an LLM make that paragraph more scientific, it is entirely your paper "More scientific" is not some merely stylistic thing that faithfully preserves the original meaning of what you wrote. The precise details of each paragraph matters a lot in terms of what and how it communicates. The fact that these details do matter means that, according to my accounting, it is not entirely y…

You might prefer that, but journal editors might not.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#58

I scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally…

this is neat! is the model available somewhere for local execution? or even a lookup table with your results for all arXiv pre-print codes? I want to run it on lots of pre-prints and I don't want to kill your server

Wouldn't that provide an excellent signal to tune models to be less like AI? I suppose that's a good thing ultimately.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#59

I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each pa…

[flagged]

Which part is a sin? Using LLMs to deal with a lack of English-language fluency? I am a scientist (actually a mathematician, if it matters), and, if that's the way to deal with the practical hegemony of English in the scientific literature, then I have no problem with it. Rather that than people with important ideas can't get them before the scientific community. As long as the authors personally check and stand behind the scientific content of their papers, what do I care how the exposition was produced, especially if the role of LLMs is properly disclosed?

Re: How we measured AI writing across arXiv, and where the measurement breaks

#60

I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each pa…

[flagged]

I feel like they should release the original before the llm chewed on it.
Post reply on HN