Live data from Hacker News

Generative AI and Wikipedia editing: What we learned in 2025

wikiedu.org

1–10 of 132 posts

Re: Generative AI and Wikipedia editing: What we learned in 2025

#2
The title I've chosen here is carefully selected to highlight one of the main points. It comes (lightly edited for length) from this paragraph:

Far more insidious, however, was something else we discovered:

More than two-thirds of these articles failed verification.

That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, the information on Wikipedia does not exist in that specific source. When a claim fails verification, it’s impossible to tell whether the information is true or not. For most of the articles Pangram flagged as written by GenAI, nearly every cited sentence in the article failed verification.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#4

The title I've chosen here is carefully selected to highlight one of the main points. It comes (lightly edited for length) from this paragraph: Far more insidious, however, was something else we discovered: More than two-thirds of these articles failed verification. That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, th…

Submitted title was "For most flagged articles, nearly every cited sentence failed verification".

I agree, that's interesting, and you've aptly expressed it in your comment here.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#5
post #3

So, a small proportion of articles were detected as bot-written, and a large proportion of those failed validation. What if in fact a large proportion of articles were bot-written, but only the unverifiable ones were bad enough to be detected?

Human editors, I suspect, would pick up the "tells" of generated text, although as we know, there's a lot of false positives in that space.

But it looks like Pangram is a text classifying NN trained using a technique where they get a human to write a body of text on a subject, and then get various LLMs to write a body of text on the same subject, which strikes me as a good way to approach the problem. Not that I'm in anyway qualified to properly understand ML.

More details here: https://arxiv.org/pdf/2402.14873

Re: Generative AI and Wikipedia editing: What we learned in 2025

#6
I feel like this is such a tragedy of the commons for the LLM providers. Wikipedia probably makes up a huge bulk of their dataset, why taint it? Would be interesting if there was some kind of "you shall not use our platform on Wikipedia" stance adopted.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#8

I feel like this is such a tragedy of the commons for the LLM providers. Wikipedia probably makes up a huge bulk of their dataset, why taint it? Would be interesting if there was some kind of "you shall not use our platform on Wikipedia" stance adopted.

I don’t think it’s the providers doing this, it’s the awful users. They’re doing the same thing on GitHub. It’s maddening.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#9

The title I've chosen here is carefully selected to highlight one of the main points. It comes (lightly edited for length) from this paragraph: Far more insidious, however, was something else we discovered: More than two-thirds of these articles failed verification. That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, th…

FWIW, this is a fairly common problem on Wikipedia in political articles, predating AI. I encourage you to give it a try and verify some citations. A lot of them turn out to be more or less bogus.

I'm not saying that AI isn't making it worse, but bad-faith editing is commonplace when it comes to hot-button topics.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#10

I feel like this is such a tragedy of the commons for the LLM providers. Wikipedia probably makes up a huge bulk of their dataset, why taint it? Would be interesting if there was some kind of "you shall not use our platform on Wikipedia" stance adopted.

It would be random individuals.
Post reply on HN