Live data from Hacker News

How we measured AI writing across arXiv, and where the measurement breaks

unslop.run

41–50 of 185 posts

Re: How we measured AI writing across arXiv, and where the measurement breaks

#41
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

The problem is that humans start with credulity. AI hallucinates and makes up stuff some percentage of the time. Humans are not "default deny" when given information. Unleashing that was a disservice to mankind and has created a future filled with lies and people who are confident in them.

LLMs are in essence an attack on the concept of written language, harvesting and dissolving the social contract that underpins it; it is no longer safe to assume that text has intent, let alone content, simply because it is grammatically well-formed.

Perhaps the most darkly amusing consequence of this particular mania is that by poisoning the majority of our information environment with hallucinated slop, we have likely crippled the next several generations of machine-learning techniques before they're even invented! Small, locally-hostable LLMs will rattle along spewing spam long after the broader "genai bubble" pops, and building clean training datasets will permanently be more difficult and expensive.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#42
post #28

Something I've thought about a lot is that there is having someone with some domain knowledge or reason to care a lot about a particular issue spend a bunch of tokens and cycles on it until something useful comes out the other end. The most obvious ones are the math problems that have been coming out and help push the frontier of various areas of math. Another example is taking all of the public NYC open data ecosyst…

Yes, I’ve thought about this too. The strength of these models is that there is a lot more knowledge encoded in them than the average scientist has in mind at any given time. That means they can explore many more possible combinations of concepts.

If we imagine a set of all human ideas that these models have access to, then the set of possible discoveries would be something like the superset of all possible combinations of those ideas. I think all LLM discoveries are bounded by that space.

Looking at the recent OpenAI math discoveries, that seems to be pretty much what happened. Existing ideas were used as building blocks, the model found a valuable combination, and the result was something new that had real value.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#43
I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors.

It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing.

If you write each paragraph and have an LLM make that paragraph more scientific, it is entirely your paper, but it is flagged as LLM generated. If you have an LLM write the paper but speak good enough English, you can make it look human even though it is not human.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#44
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

you just dont respect your readers, thats all

if 5% of the work that I'm reading to keep up is slop, that's a shame. that's time I spent puzzling over connections that weren't real, trying to impose some kind of logic on the arguments, and wondering why the data doesn't really support the conclusion.

if 50% of the work is nonsense, then there's a serious concern that we can't move forward at all.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#45
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

Another important questions is:

What about the papers that graduate to proper publication?

Arxiv is full of pre-prints that anyone can upload.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#47
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

I recently desk-rejected a paper where every single citation in its Introduction was hallucinated. That means that the entire connection between what the author(s) did and how it relates to existing research was simply made up. I've never seen this happening before AI but now there's at least one paper in every cycle pulling something similar. My problem therefore is: we are seeing more and more papers written with t…

While working on my PhD, there was one guy in the field that would publish new, better results to Springer every time someone improved upon benchmarks. No code, no data, no conference publications, paywall restricted, unverifiable, always the best.

I am almost certain he was "hallucinating" the results. This was in the 2010s

There are well known issues in academic publishing, though I imagine it has become much noisier like open source

Re: How we measured AI writing across arXiv, and where the measurement breaks

#49

I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each pa…

There is a fine line between using an LLM to clean up grammar and spelling and using it for text generation.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#50

I am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each pa…

[flagged]
Post reply on HN