Live data from Hacker News

Wikipedia: WikiProject AI Cleanup

en.wikipedia.org

81–90 of 99 posts

Re: Wikipedia: WikiProject AI Cleanup

#81
post #56

Earlier quoted context omitted.

You can easily do this with normal GPT 5.2 in ChatGPT, just turn on thinking (better if extended) and web search, point a Wikipedia page to the model and tell it to check the claims for errors. I've tried it before and surprisingly it finds errors very often, sometimes small, sometimes medium. The less popular the page you linked is, the more likely it'll have errors. This works because GPT 5.x actually properly use…

It says it finds errors.

It gives references that you can then verify manually. I wasn't advocating for a 100% automated process.

Re: Wikipedia: WikiProject AI Cleanup

#82

There was a paper recently about using LLMs to find contradictions in Wikipedia, i.e. claims on the same page or between pages which appear to be mutually incompatible. https://arxiv.org/abs/2509.23233 I wonder if something more came out of that. Either way, I think that generation of article text is the least useful and interesting way to use AI on Wikipedia. It's much better to do things like this paper did.

You can easily do this with normal GPT 5.2 in ChatGPT, just turn on thinking (better if extended) and web search, point a Wikipedia page to the model and tell it to check the claims for errors. I've tried it before and surprisingly it finds errors very often, sometimes small, sometimes medium. The less popular the page you linked is, the more likely it'll have errors. This works because GPT 5.x actually properly use…

Can you describe some of the errors you have found this way?

Re: Wikipedia: WikiProject AI Cleanup

#83
post #82

Earlier quoted context omitted.

You can easily do this with normal GPT 5.2 in ChatGPT, just turn on thinking (better if extended) and web search, point a Wikipedia page to the model and tell it to check the claims for errors. I've tried it before and surprisingly it finds errors very often, sometimes small, sometimes medium. The less popular the page you linked is, the more likely it'll have errors. This works because GPT 5.x actually properly use…

Can you describe some of the errors you have found this way?

Usually they are smaller details in the pages, not the core claims, but that doesn't really refute my point of GPT being easily able to point them. Here are some examples, I'm not including all errors per page that GPT 5.1 found back then (this reply is already too long), but just a few examples.

https://en.wikipedia.org/wiki/Large_Hadron_Collider

> (infobox) Maximum luminosity 1×10^34/(cm2⋅s)

This is from the original design, LHC has been upgraded several times, e.g. if you check https://home.web.cern.ch/news/news/accelerators/lhc-report-r..., you see "Thanks to these improvements, the instantaneous luminosity record was smashed, reaching 2.06 x 10^34 cm^(-2) s^(-1), twice the nominal value." and that was in 2017.

> The first collisions were achieved in 2010 at an energy of 3.5 tera-electronvolts (TeV) per beam

This is wrong, if you check https://home.web.cern.ch/resources/faqs/facts-and-figures-ab... it says "23 November 2009: LHC first collisions (see press release)" - https://home.web.cern.ch/news/press-release/cern/two-circula... and the energy was 450 GeV

Another random example, I was reading https://en.wikipedia.org/wiki/Camponotus_japonicus (a very small article) and decided to ask GPT about it. It checked a lot of other sources and found out that no other source claims that this species of ant inhabits Iran.

Another one: https://en.wikipedia.org/wiki/Java_(software_platform)

> and—until its discontinuation in JDK 9—a browser plug-in

In reality it was deprecated in JDK 9 and removed in JDK 11 - most people would think "discontinuation" means that it was already removed in JDK 9

https://en.wikipedia.org/wiki/Nekopara

> The Opening theme for After, "Contrail" was composed by "Motokyio" and Sung by "Ceul".

Just two misspellings, it should be Motokiyo and Ceui

> A manga adaptation illustrated by Tam-U is currently being published

This section hasn't been updated, but the manga has already finished a long time ago.

===

Here's a direct part of GPT 5.1's response (I tried this back in November, so there was no GPT 5.2 yet) regarding luminosity, and it did also have a citation in the 2nd paragraph to the exact link I used above for the luminosity claim.

– The infobox lists “Maximum luminosity 1×10^34/(cm²·s)” without qualification.

– That number is the original design (nominal) peak luminosity for the LHC, but the machine has substantially exceeded it in routine operation: CERN operations reports show peak instantaneous luminosities of about 1.6×10^34 cm⁻²·s⁻¹ in 2016 and ≈2.0–2.1×10^34 cm⁻²·s⁻¹ in 2017–2018, roughly a factor of two above the nominal design.

– Since the same infobox uses the current maximum beam energy (6.8 TeV per beam) rather than the 7 TeV design value, presenting 1×10^34 cm⁻²·s⁻¹ as “Maximum luminosity” is misleading/outdated if read as the machine’s achieved maximum. It should either be labelled explicitly as “design luminosity” (with a note that higher values have been reached) or the numerical value should be updated to reflect the achieved peak.

Re: Wikipedia: WikiProject AI Cleanup

#84

Contrarian take: Wikipedia could use more AI, as well as less. A major flaw of Wikipedia is that much of it is simply poorly written. Repetition and redundancy, ambiguity, illogical ordering of content, rambling sentences, opaque grammar. That should not be surprising. Writing clear prose is a skill that most people do not have, and Wikipedia articles are generally the fruit of collaboration without copy editors. AI…

> AI is perfectly suited to fixing this problem. I recently spent several hours rewriting a somewhat important article. I did not add or subtract information from the article, I simply made it clearer and more concise. I'm confused by this. Is this written by an AI? > Repetition and redundancy, ambiguity, illogical ordering of content, rambling sentences, opaque grammar. This pile of words is missing a verb. "You" (w…

> This pile of words is missing a verb.

And yet is completely understandable.

Re: Wikipedia: WikiProject AI Cleanup

#85
post #75

I wish they also spent on the reverse: automatic rephrasing of the (many) obscure and very poorly worded and/or with no neutral tone whatsoever. And I say that as a general Wikipedia fan.

there are many copy editing projects that do this. If you mean the left leaning tone / bias, that will be a bit more spicy. But general grammar, tone, ambiguity , superlatives – that’s the goal of copy editing. I copy edit typesetting , for example.

> If you mean the left leaning tone / bias, that will be a bit more spicy. But general grammar, tone, ambiguity , superlatives – that’s the goal of copy editing

No, no I mainly mean non-neutral phrasing and/or too personal. Especially for people’s articles. (“And they released that greeeat album! But unfortunately the critics did not understand them… Booh!)

Re: Wikipedia: WikiProject AI Cleanup

#86

Contrarian take: Wikipedia could use more AI, as well as less. A major flaw of Wikipedia is that much of it is simply poorly written. Repetition and redundancy, ambiguity, illogical ordering of content, rambling sentences, opaque grammar. That should not be surprising. Writing clear prose is a skill that most people do not have, and Wikipedia articles are generally the fruit of collaboration without copy editors. AI…

> AI is perfectly suited to fixing this problem. I recently spent several hours rewriting a somewhat important article. I did not add or subtract information from the article, I simply made it clearer and more concise. I'm confused by this. Is this written by an AI? > Repetition and redundancy, ambiguity, illogical ordering of content, rambling sentences, opaque grammar. This pile of words is missing a verb. "You" (w…

Hard to explain the hostility here. I simply outlined my opinion ("take") and backed it up with reasons. I have been a Wikipedia editor for well over 20 years. That should not be relevant to my argument.

Re: Wikipedia: WikiProject AI Cleanup

#87

I opened a random page with the label: https://en.wikipedia.org/wiki/Ain%27t_in_It_for_My_Health Curious, what are the signs that this particular page has been written by an AI? I’m not saying it wasn’t, I’m probably not seeing something and wondering what to look for.

Likely this passage:

>Upon release, the album received generally positive reviews from critics, with praise for Top's traditionalist approach and vocal authenticity, though some noted its adherence to familiar country frameworks.

Generic and uncited.

Re: Wikipedia: WikiProject AI Cleanup

#88
post #11

Earlier quoted context omitted.

An interesting observation from that page: "Thus the highly specific "inventor of the first train-coupling device" might become "a revolutionary titan of industry." It is like shouting louder and louder that a portrait shows a uniquely important person, while the portrait itself is fading from a sharp photograph into a blurry, generic sketch. The subject becomes simultaneously less specific and more exaggerated."

That sounds like Flanderization to me https://en.wikipedia.org/wiki/Flanderization From my experience with LLMs that's a great observation.

[dead]

Re: Wikipedia: WikiProject AI Cleanup

#89
post #11

Earlier quoted context omitted.

An interesting observation from that page: "Thus the highly specific "inventor of the first train-coupling device" might become "a revolutionary titan of industry." It is like shouting louder and louder that a portrait shows a uniquely important person, while the portrait itself is fading from a sharp photograph into a blurry, generic sketch. The subject becomes simultaneously less specific and more exaggerated."

I think that's a general guideline to identify "propaganda", regardless of the source. I've seen people in person write such statements with their own hands/fingers, and I know many people who speak like that (shockingly, most of them are in management). Lots of those points seems to get into the same idea which seems like a good balance. It's the language itself that is problematic, not how the text itself came to b…

A good place to see this pre-2022 (the ai epoch) is articles on less known bands from the late 2000s when Wikipedia was becoming more popular. Quite a few of them turn out to be copy/paste promo text. I know this because I did webdev work for that industry, and when I look up those bands on wikipedia I will recognize the text as text that I personally had to paste into a bio page 20 years ago. Since the bands are well known, nobody reports it (I admit I'm too lazy)

The real tell on those tends to be weirdly time-specific claims that tend to be wildly outdated ("currently touring with XYZ")

Re: Wikipedia: WikiProject AI Cleanup

#90
post #75

Earlier quoted context omitted.

there are many copy editing projects that do this. If you mean the left leaning tone / bias, that will be a bit more spicy. But general grammar, tone, ambiguity , superlatives – that’s the goal of copy editing. I copy edit typesetting , for example.

> If you mean the left leaning tone / bias, that will be a bit more spicy. But general grammar, tone, ambiguity , superlatives – that’s the goal of copy editing No, no I mainly mean non-neutral phrasing and/or too personal. Especially for people’s articles. (“And they released that greeeat album! But unfortunately the critics did not understand them… Booh!)

I agree. Wikipedia Cleanup is a good starting point. Or look for a Wiki Project to join.

I've found the best way to learn and contribute is to jump into an existing project. Usually direction is the hardest thing .

You can of course dive into an article and make changes, but you'll often get pushback (warranted or unwarranted) and that can be discouraging. It's a somewhat natural feedback loop.

Post reply on HN