Live data from Hacker News

How we measured AI writing across arXiv, and where the measurement breaks

unslop.run

81–90 of 185 posts

Re: How we measured AI writing across arXiv, and where the measurement breaks

#81
I’m very skeptical of these results. I run a pretty large group paper website on top of arXiv and we run pangram on papers. The numbers are not nearly this high.

One thing I see a lot is papers flagged as AI because they include llm rollouts in the paper as examples.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#82
I run most professional statements and articles through LLMs before submission. This helps correct grammar and improves accessibility through better syntax (because I write exactly how I think).

The article doesn't seem to mention consideration of AI for polishing human work.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#83

I run most professional statements and articles through LLMs before submission. This helps correct grammar and improves accessibility through better syntax (because I write exactly how I think). The article doesn't seem to mention consideration of AI for polishing human work.

> I run most professional statements and articles through LLMs before submission. This helps correct grammar and improves accessibility through better syntax (because I write exactly how I think).

> The article doesn't seem to mention consideration of AI for polishing human work.

Because it isn't a consideration. You are what they are looking for.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#84
post #65

I'm not sure what to think... I uploaded a PyHPC workshop paper I wrote in 2011 and it said 27% machine. I also uploaded my PhD dissertation from 2012 and got back 40% machine, which is just barely below the 42% threshold. I don't publish anymore... but does this mean I wrote like a LLM or did LLMs learn from me? :p Update: I also uploaded a IEEE CLUSTERS paper I wrote in 2015 and it came back 74% machine written :|

When "detectors" first started popping up all over the place, all of them rated the declaration of independence as 100% AI written, so... Yeah, these things just don't work. And what's even more dangerous is that people that don't understand how any of it works use these tools, and accuse people of using AI, sometimes with grave consequences. Students have been through this, at all levels of education.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#85
There some real game theory mechanics at play when it comes to LLM usage in corporations. Devs are cranking out superficially superior code and documentation by just aiming the Claude Code fire hose at everything they can. Leadership encourages this because from what they can tell, there is no downside.

It's hard to say if this code is structurally better or worse than before, but it's certainly voluminous and as far as leadership could ever tell, with their flawed metrics, that is all that matters. It will be years before we figure out if this is a good idea and worth the cognitive atrophy.

Anyone not using LLMs all day is just not going to be as prolific. I can't imagine that the same factors aren't at play in the scientific research community where it's all about how much you can publish.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#86
post #28

Something I've thought about a lot is that there is having someone with some domain knowledge or reason to care a lot about a particular issue spend a bunch of tokens and cycles on it until something useful comes out the other end. The most obvious ones are the math problems that have been coming out and help push the frontier of various areas of math. Another example is taking all of the public NYC open data ecosyst…

Yes, I’ve thought about this too. The strength of these models is that there is a lot more knowledge encoded in them than the average scientist has in mind at any given time. That means they can explore many more possible combinations of concepts. If we imagine a set of all human ideas that these models have access to, then the set of possible discoveries would be something like the superset of all possible combinati…

Yeah and I think even what might be the most useful is extracting how Codex got to a particular solution into a skill or into even more customized software so it can be applied in lots more places.

I know people are working on these things I just haven’t seen the right way yet. Like in manufacturing right now people are trying to encode what skilled machinists do into software and scale it up, we need to go further on that for Math/ data analysis etc.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#87

I run most professional statements and articles through LLMs before submission. This helps correct grammar and improves accessibility through better syntax (because I write exactly how I think). The article doesn't seem to mention consideration of AI for polishing human work.

AI doesn't polish human work. It is more like an extruder that squeezes anything you put into it into a generic, formulaic shape, indistinguishable from writing that was lazily prompted because the writer couldn't be bothered to put forth any more effort than that.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#88
post #85

There some real game theory mechanics at play when it comes to LLM usage in corporations. Devs are cranking out superficially superior code and documentation by just aiming the Claude Code fire hose at everything they can. Leadership encourages this because from what they can tell, there is no downside. It's hard to say if this code is structurally better or worse than before, but it's certainly voluminous and as far…

> Leadership encourages this because from what they can tell, there is no downside.

This was my big fear before we saw price increases. Now I'm pinning all my hopes on AI being too expensive to justify further big corporate pushes. (Sigh.) I love having new tools, but I hate being pushed to use ______ tool to meet some managerial metric.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#89

Earlier quoted context omitted.

I recently desk-rejected a paper where every single citation in its Introduction was hallucinated. That means that the entire connection between what the author(s) did and how it relates to existing research was simply made up. I've never seen this happening before AI but now there's at least one paper in every cycle pulling something similar. My problem therefore is: we are seeing more and more papers written with t…

Yeah but that's an issue with the researcher putting out a bad paper, and it suggests you'll have to reject more papers. We wouldn't ban email because many of the emails are spam, it just means we need new tools to filter out junk. AI will allow researchers to be more productive all together and take less time to publish a paper, which is good.

> Yeah but that's an issue with the researcher putting out a bad paper, and it suggests you'll have to reject more papers. We wouldn't ban email because many of the emails are spam, it just means we need new tools to filter out junk. AI will allow researchers to be more productive all together and take less time to publish a paper, which is good.

It's a signal:noise ratio thing. If 1 out of every 1000 AI-written papers are bad, it makes sense to put in a filter that auto-rejects any paper that has AI tells.

After all, if that 1 researcher was any good, he wouldn't have used AI to write the thing in the first place.

Publishing was always about getting past the filters. There's one more filter - "AI-generated content" - so do what you have to to get past it. IOW, write your own paper.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#90

I scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally…

Extend it to scan papers from pre-2020. That should give you a better baseline accuracy for your detection system.
Post reply on HN