Live data from Hacker News

Generative AI and Wikipedia editing: What we learned in 2025

wikiedu.org

81–90 of 132 posts

Re: Generative AI and Wikipedia editing: What we learned in 2025

#81
post #41

Earlier quoted context omitted.

There was a fun example of this that happened live during a recent episode of the Changelog[1]. The hosts noted that they were incorrectly described as being "from GitHub" with a link to an episode of their podcast which didn't substantiate that claim. Their guest fixed the citation as they recorded[2]. [1]: https://changelog.com/podcast/668#transcript-265 [2]: https://en.wikipedia.org/w/index.php?title=Eugen_Rochko&…

How did they know it was not LLM generated?

The false claim was added 7 Nov 2022 [1] while chatgpt wasn't released until 30 Nov 2022.

[1] https://en.wikipedia.org/w/index.php?title=Eugen_Rochko&diff...

Re: Generative AI and Wikipedia editing: What we learned in 2025

#82
post #36

Earlier quoted context omitted.

LLMs can add unsubstantiated conclusions at a far higher rate than humans working without LLMs.

True, but humans got a 20 year head start and I am willing to wager the overwhelming majority of extant flagrant errors are due to humans making shit up and no other human noticing and correcting it. My go too example was the SDI page saying that brilliant pebble interceptors were to be made out of tungsten (completely illogical hogwash that doesn't even pass a basic sniff test.) This claim was added to the page in F…

If LLMs 10X it, as the advocates keep insisting, that means it would only take 2 years to do as much or more damage as humans alone have done in 20.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#83

I feel like this is such a tragedy of the commons for the LLM providers. Wikipedia probably makes up a huge bulk of their dataset, why taint it? Would be interesting if there was some kind of "you shall not use our platform on Wikipedia" stance adopted.

I don’t think it’s the providers doing this, it’s the awful users. They’re doing the same thing on GitHub. It’s maddening.

I don't think Lockheed Martin or Raytheon are doing this, it's the awful pilots and intercept operators launching missiles into Palestinian homes. I don't think Rostec Corporation is doing this. It's only the grunts on the ground pressing the button sending heavy munitions into crowds of Ukranian civilians.

These mega corporations are entirely free from blame and you're gonna see to it none of us question their role, right?

Re: Generative AI and Wikipedia editing: What we learned in 2025

#84
So, AI spam can degrade quality.

But ... isn't this with regards to Wikipedia a much more general problem?

Usually revisions are approved manually by real people. This already can be negative; takes a lot of time; no guarantee that new information is true but old information can be wrong too. To me it seems more as if the problem has much more to do with the quality control problems of wikipedia itself. Yes, AI spam fatigues here but if the quality control steps are bad then AI spam will only make this worse. But AI spam going away, does not mean the quality control steps have gotten any better. These two issues should be separate. Wikipedia needs to find better quality control mechanisms in general. And that also includes existing articles - some are written by people who are experts in the field. But they don't really explain anything at all. So, these articles appear good but are virtually useless for 98% of the people. I am not saying one should dumb down wikipedia, but you need to kind of focus primarily on the average person really - not stupid but not a godlike expert either. Explain it to, say, someone at age 18 or perhaps even a bit less than that.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#85

Earlier quoted context omitted.

True, but humans got a 20 year head start and I am willing to wager the overwhelming majority of extant flagrant errors are due to humans making shit up and no other human noticing and correcting it. My go too example was the SDI page saying that brilliant pebble interceptors were to be made out of tungsten (completely illogical hogwash that doesn't even pass a basic sniff test.) This claim was added to the page in F…

If LLMs 10X it, as the advocates keep insisting, that means it would only take 2 years to do as much or more damage as humans alone have done in 20.

Perhaps so. On the other hand, there's probably a lot of low hanging fruit they can pick just by reading the article, reading the cited sources, and making corrections. Humans can do this, but rarely do because it's so tedious.

I don't know how it will turn out. I don't have very high hopes, but I'm not certain it will all get worse either.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#86

> That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, the information on Wikipedia does not exist in that specific source. When a claim fails verification, it’s impossible to tell whether the information is true or not. This has been a rampant problem on Wikipedia always. I can't seem to find any indicator that this has…

> Applying correct citations is actually really hard work

Not disagreeing - many existing articles on wikipedia have barely any references or citation at all and in some cases wrong citation or wrong conclusions. Like when an article says water molecules behave oddly and then the wikipedia article concluding that water molecules behave properly.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#87

Earlier quoted context omitted.

It’s on its way to becoming more popular and a clear competitor to it. Just a matter of time.

Is that "more popular" in the sense of McDonald's popular?

Yeah so?

Re: Generative AI and Wikipedia editing: What we learned in 2025

#89
post #58

Earlier quoted context omitted.

The problems I've run into is both people giving fake citations (the citations don't actually justify the claim that's being made in the article), and people giving real citations, but if you dig into the source you realize it's coming from a crank. It's a big blind spot among the editors as well. When this problem was brought up here in the past, with people saying that claims on Wikipedia shouldn't be believed unle…

> but if you dig into the source you realize it's coming from a crank. It is a dark sunday afternoon, Bob Park is sitting on his sofa as usual, drunk as usual, suddenly the TV reveals to him there to be something called the Paranormal (Twilight Zone music) ..instantly Bob knows there are no such things and adds a note to the incomprehensible mess of notes that one day will become his book. He downs one more Budweiser…

Curious what the point you're making here is. I don't know anything at all about Bob Park and whether he is a crank. But if you make your career doing the admirable work of debunking pseudo-science and nonsense theories, you would necessarily be linked to in discussions of those theories very, very frequently.

So maybe that's not a good description of him. But the link you posted is hardly dispositive.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#90

Earlier quoted context omitted.

The problems I've run into is both people giving fake citations (the citations don't actually justify the claim that's being made in the article), and people giving real citations, but if you dig into the source you realize it's coming from a crank. It's a big blind spot among the editors as well. When this problem was brought up here in the past, with people saying that claims on Wikipedia shouldn't be believed unle…

A common source of error is in articles for movies where it gives plot summaries. The plot summaries are very often written by people who didn't watch the movie but are trying to re-resemble the plot like a jigsaw puzzle from little bits they glean from written reviews, or worse just writing down whatever they assume to be the plot. Very often it seems like the fuck ups came from people who either weren't watching th…

To me that sounds a bit like summaries made on the base of written movie scripts. A long time ago, I read a few scripts to movies I had never watched, and that's exactly the outcome: You get a rough idea what it's about and even get to recognise some memorable quotes, but there's little cohesion to it, for lack of all the important visual aspects and clues that tie it all together.
Post reply on HN