Live data from Hacker News

Generative AI and Wikipedia editing: What we learned in 2025

wikiedu.org

111–120 of 132 posts

Re: Generative AI and Wikipedia editing: What we learned in 2025

#111
post #17

> That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, the information on Wikipedia does not exist in that specific source. When a claim fails verification, it’s impossible to tell whether the information is true or not. This has been a rampant problem on Wikipedia always. I can't seem to find any indicator that this has…

When I've checked Wikipedia citations I've found so much brazen deception - citations that obviously don't support the claim - that I don't have confidence in Wikipedia. > Applying correct citations is actually really hard work, even when you know the material thoroughly. Why do you find it hard? Scholarly references can be sources for fundamental claims, review articles are a big help too. Also, I tend to add things…

I found several that were contradicting the claim they were supposed to support (in popular articles). I will never regain faith in wikipedia. Being an editor or just verifying information from wikipedia makes you hate it

Re: Generative AI and Wikipedia editing: What we learned in 2025

#112

Earlier quoted context omitted.

If LLMs 10X it, as the advocates keep insisting, that means it would only take 2 years to do as much or more damage as humans alone have done in 20.

Perhaps so. On the other hand, there's probably a lot of low hanging fruit they can pick just by reading the article, reading the cited sources, and making corrections. Humans can do this, but rarely do because it's so tedious. I don't know how it will turn out. I don't have very high hopes, but I'm not certain it will all get worse either.

The entire point of the article is that LLMs cannot make accurate text, but ironically you claiming LLMs can do accurate texts illustrates your point about human reliability perfectly.

I guess the conclusion is there simply is no avenues to gain knowledge.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#113

> That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, the information on Wikipedia does not exist in that specific source. When a claim fails verification, it’s impossible to tell whether the information is true or not. This has been a rampant problem on Wikipedia always. I can't seem to find any indicator that this has…

> This has been a rampant problem on Wikipedia always. I can't seem to find any indicator that this has increased recently? Because they're only even investigating articles flagged as potentially AI. So what's the control baseline rate here? ...y'know, I don't want to be that guy, but this actually seems like something AI could check for, and then flag for human review.

You don't need an LLM to find loads of uncorroborated claims on Wikipedia. See f.e. https://en.wikipedia.org/wiki/Variational_autoencoder Most articles about tech are woefully undersourced.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#114
post #91

Earlier quoted context omitted.

It's not that I personally find it hard. It's more like, a lot of stuff in Wikipedia articles is somewhat "general" knowledge in a given field, where it's not always exactly obvious how to cite it, because it's not something any specific person gets credit for "inventing". Like, if there's a particular theorem then sure you cite who came up with it, or the main graduate-level textbook it's taught in. But often it's j…

If you can buy the food at a supermarket, can't you cite a product page? Presumably that would include a description of the product. Or is that not good enough of a citation?

Retail product listing URLs change constantly. They're not great.

And then you usually want to describe how the food is used. E.g. suppose it's a dessert that's mainly popular at children's birthday parties. Everybody in the country knows that. But where are you going to find something written that says that? Something that's not just a random personal blog, but an actual published valid source?

Ideally you can find some kind of travel guide or book for expats or something with a food section that happens to list it, but if it's not a "top" food highly visible to tourists, then good luck.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#115
post #57

I’m honestly surprised LLMs are still screwing up citations. It does not feel like a harder task than building software or generating novel math proofs. In both those cases, of course, there is a verifier, but self-verification with “Does this text support this claim?” seems like it ought to be within the capabilities of a good reasoning model. But as I understand the situation, even the major Deep Research systems s…

> LLMs [...] reasoning model

Found your problem right there

Re: Generative AI and Wikipedia editing: What we learned in 2025

#116
post #109

Earlier quoted context omitted.

Is that "more popular" in the sense of McDonald's popular?

Doesnt really matter the reason why does it? If more people use it then it is more popular. Simple metric.

Well, kinda ...

Approximately 3.2m Americans drink Hard Kombucha.

Approximately 5m Americans use Cocaine.

Simple metric.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#119
post #36

Earlier quoted context omitted.

LLMs can add unsubstantiated conclusions at a far higher rate than humans working without LLMs.

True, but humans got a 20 year head start and I am willing to wager the overwhelming majority of extant flagrant errors are due to humans making shit up and no other human noticing and correcting it. My go too example was the SDI page saying that brilliant pebble interceptors were to be made out of tungsten (completely illogical hogwash that doesn't even pass a basic sniff test.) This claim was added to the page in F…

> I am willing to wager the overwhelming majority of extant flagrant errors are due to humans making shit up

In general, I agree, but I wouldn't want to ascribe malfeasance ("making shit up") as the dominant problem.

I've seen two types of problems with references.

1. The reference is dead, which means I can't verify or refute the statement in the Wikipedia article. If I see that, I simply remove both the assertion and the reference from the wiki article.

2. The reference is live, but it almost confirms the statement in the wikipedia article, but whoever put it there over-interpreted the information in the reference. In that case, I correct the statement in the article, but I keep the ref.

Those are the two types of reference errors that I've come across.

And, yes, I've come across these types of errors long before LLMs.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#120
post #25

Earlier quoted context omitted.

This is like saying handing out machine guns is no big change because people have been shooting arrows for a long time. At some point volume becomes the story once it overwhelms the community’s ability to correct errors.

> once it overwhelms the community’s ability to correct errors. I think the point is that it already has.

I wouldn’t be so quick to write Wikipedia off but I do think there will be changes. The SEO abusers might not have cost us editor anonymity but LLMs might push in the direction of needing real-world validation or friend of a friend referrals.
Post reply on HN