Live data from Hacker News

How we measured AI writing across arXiv, and where the measurement breaks

unslop.run

31–40 of 185 posts

Re: How we measured AI writing across arXiv, and where the measurement breaks

#31
post #11

Earlier quoted context omitted.

The problem is that humans start with credulity. AI hallucinates and makes up stuff some percentage of the time. Humans are not "default deny" when given information. Unleashing that was a disservice to mankind and has created a future filled with lies and people who are confident in them.

Oh, to be technically correct: AI hallucinates and makes up stuff 100% percent of the time. Never been a fan of that word for this. Again, I fail to see the problem here that isn't solved by careful reading WHICH IS WHAT PEOPLE SHOULD BE DOING ANYWAY. I would like to see room for AI disclosure, maybe a statement of "this is how much AI I used." But this blanket X% of this looks like AI? Again, so what?

I find the fact that 39% of papers use a style that signals lack of effort somewhat worrisome. Even if that’s 100% wrong, and they’re all high effort papers, the fact that they give off the same aura as low effort papers is a problem for the authors.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#32
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

Surprisingly I agree with you. My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

I reviewed for ACL (big NLP conf) and another similar conference recently. Lots of crap (though there always has been). One superficially well-written paper was likely AI-plagiarized (i.e. the AI sampled an idea from prior work and rewrote it) and got desk rejected for it.

Reviewers were also totally unengaged. Of 20 reviews I read (from my reviewers or from reviewers on the same papers), maybe 2 were mediocre, and the rest were crap (though likely not AI).

The notion that science will somehow benefit from this is about as stupid an idea as you can have. Science relies on skepticism. AIs are not skeptical, and many folks are submitting papers because they stand to gain something, not because they are motivated to do good research or develop new understanding. Fields are being inundated with garbage that is maximally indistinguishable from real work (that's the training objective for LLMs). This in turn maximizes the cost of identifying bad work.

This is the same enshittification process that we see everywhere else. You get spam phone calls because there is no reason for a spammer not to call you. "Researchers" are submitting spam papers because there is no cost to doing so with some possible gain. Absent intervention, this eventually drives the community value of the network to zero (or potentially negative, if friction costs to switching are high).

Re: How we measured AI writing across arXiv, and where the measurement breaks

#34
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

Surprisingly I agree with you. My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

I think the key question is more effective at what?

I see plenty of anecdotal evidence that models have been trained fantastically well—and getting better—at writing to trigger the right neurons in the human population to produce “This is interesting/informative/correct” responses in bulk.

Could their ability to produce those responses run far ahead of their ability to actually achieve the last in reality? Sure seems plausible, and then where are we?

Re: How we measured AI writing across arXiv, and where the measurement breaks

#35
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

Surprisingly I agree with you. My opinion is: if it makes communicating research more effective, while not reducing the quality of the output substantially, I see no issue. A possible conclusion for this could be: If the majority of CS papers is AI written, let's just accept this reality universally and stop worrying about it altogether.

> if it makes communicating research more effective, while not reducing the quality of the output substantially

That’s a big if. We all know that’s not what’s happening.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#36
post #11

Earlier quoted context omitted.

The problem is that humans start with credulity. AI hallucinates and makes up stuff some percentage of the time. Humans are not "default deny" when given information. Unleashing that was a disservice to mankind and has created a future filled with lies and people who are confident in them.

Oh, to be technically correct: AI hallucinates and makes up stuff 100% percent of the time. Never been a fan of that word for this. Again, I fail to see the problem here that isn't solved by careful reading WHICH IS WHAT PEOPLE SHOULD BE DOING ANYWAY. I would like to see room for AI disclosure, maybe a statement of "this is how much AI I used." But this blanket X% of this looks like AI? Again, so what?

If as a primary school student you had needed to reverify every fact /experiment that was presented in your textbooks, you would not have gotten very far.

Trust is very important to human progress.

Re: How we measured AI writing across arXiv, and where the measurement breaks

#38
the funniest part of these AI detectors is that if I were to upload any of einstin's paper's they will all be flagged as AI-written. it makes sense because it's part of their training data.

but this post makes me wonder, if more papers' are written with AI, or the shape of knowledge of converging?

Re: How we measured AI writing across arXiv, and where the measurement breaks

#39
post #3

The important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describ…

LLM-written text tends to have the property that it gives the impression of expertise and knowledge in excess of what the text actually contains.

In other words, it "sounds smart" without necessarily having anything to back it up. In even more critical terms, it's very good at bullshitting.

Unfortunately for us, the scientific community current relies on a certain amount of trust. (To do otherwise is very expensive! see: bitcoin). When you introduce a known-bullshitter to write your papers, every human in the loop effectively has to defend against an adversarial attack. Not just the readers at home, or the peer reviewers, but even the author needs to be wary that the facts and arguments coming out of the LLM are true and meaningful.

Personally, I've been a minor contributor to several high-profile papers. I don't know how every field does it, but in my experience, the corresponding author (generally the PI or other senior scientist), is responsible for the accuracy of the paper. They ultimately have to trust the people who did the work that the facts are true. Introducing LLMs into the mix make it more difficult for them to identify and review sections they are unsure of. (An honest person will typically write at a confidence level reflecting their certainty. LLMs do not do this in any reliable way.)

I've also found that LLMs frequently use metaphors that are unhelpful, or used out of context in a field that isn't familiar with them. This makes understanding the text more work, for no good reason. Introducing terms or definitions with low relevance reads as impressive at first glance, but avoiding the standard terminology in the field just adds confusion. (As an analogy, imagine if you were reading a CS paper that, for no particular reason, devoted a section to a new data structure called an "akimbo tree," which after much untangling, you realized was a reinvention of a randomized splay tree.)

Re: How we measured AI writing across arXiv, and where the measurement breaks

#40
As with all text-only AI detection schemes, I am concerned about the accuracy of the detection. I'm skeptical of the methodology, specifically the final join of the three detector scores---how can you be sure that that final step does not introduce any biases? There's no source available, so it's difficult to tell exactly how this works or reproduce the research.

I've also uploaded text samples from my own (unreleased) research from pre-LLM era, and it's seemingly scoring pretty high on the LLM-detection scores. On other papers, nearly every sentence is highlighted as red "machine-leaning," but that does not impact the score? Additionally, there are dramatic differences between the scores for identical text with and without LaTeX formatting, despite the fact that it should not matter.

The takeaway from this should be "it is difficult to detect generated text and we should be careful about accepting results simply because they confirm a hypothesis."

--

Relatedly, the text above scores as highly machine-written, despite the fact that I just wrote it with my human hands, I promise :)

Post reply on HN