Generative AI and Wikipedia editing: What we learned in 2025
1–10 of 132 posts
Re: Generative AI and Wikipedia editing: What we learned in 2025
#2Far more insidious, however, was something else we discovered:
More than two-thirds of these articles failed verification.
That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, the information on Wikipedia does not exist in that specific source. When a claim fails verification, it’s impossible to tell whether the information is true or not. For most of the articles Pangram flagged as written by GenAI, nearly every cited sentence in the article failed verification.
Re: Generative AI and Wikipedia editing: What we learned in 2025
#3What if in fact a large proportion of articles were bot-written, but only the unverifiable ones were bad enough to be detected?
Re: Generative AI and Wikipedia editing: What we learned in 2025
#4The title I've chosen here is carefully selected to highlight one of the main points. It comes (lightly edited for length) from this paragraph: Far more insidious, however, was something else we discovered: More than two-thirds of these articles failed verification. That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, th…
I agree, that's interesting, and you've aptly expressed it in your comment here.
Re: Generative AI and Wikipedia editing: What we learned in 2025
#5So, a small proportion of articles were detected as bot-written, and a large proportion of those failed validation. What if in fact a large proportion of articles were bot-written, but only the unverifiable ones were bad enough to be detected?
But it looks like Pangram is a text classifying NN trained using a technique where they get a human to write a body of text on a subject, and then get various LLMs to write a body of text on the same subject, which strikes me as a good way to approach the problem. Not that I'm in anyway qualified to properly understand ML.
More details here: https://arxiv.org/pdf/2402.14873
Re: Generative AI and Wikipedia editing: What we learned in 2025
#6Re: Generative AI and Wikipedia editing: What we learned in 2025
#7Re: Generative AI and Wikipedia editing: What we learned in 2025
#8I feel like this is such a tragedy of the commons for the LLM providers. Wikipedia probably makes up a huge bulk of their dataset, why taint it? Would be interesting if there was some kind of "you shall not use our platform on Wikipedia" stance adopted.
Re: Generative AI and Wikipedia editing: What we learned in 2025
#9The title I've chosen here is carefully selected to highlight one of the main points. It comes (lightly edited for length) from this paragraph: Far more insidious, however, was something else we discovered: More than two-thirds of these articles failed verification. That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, th…
I'm not saying that AI isn't making it worse, but bad-faith editing is commonplace when it comes to hot-button topics.
Re: Generative AI and Wikipedia editing: What we learned in 2025
#10I feel like this is such a tragedy of the commons for the LLM providers. Wikipedia probably makes up a huge bulk of their dataset, why taint it? Would be interesting if there was some kind of "you shall not use our platform on Wikipedia" stance adopted.