Live data from Hacker News

Generative AI and Wikipedia editing: What we learned in 2025

wikiedu.org

121–130 of 132 posts

Re: Generative AI and Wikipedia editing: What we learned in 2025

#121

Earlier quoted context omitted.

Is the sarcasm really that opaque? Who would unironically equate buggy accidents and automobile accidents?

How much time have you spent around developers?

I got my first tech job in 1998. Some of the most sarcastic people I’ve ever met.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#122
post #55

Earlier quoted context omitted.

Is the sarcasm really that opaque? Who would unironically equate buggy accidents and automobile accidents?

I’d like to introduce you to the internet. There’s a reason /s was a big thing, one persons obvious sarcasm is (almost tautologically) another persons true statement of opinion.

Thanks. I wasn’t aware of that.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#123
post #41

Earlier quoted context omitted.

How did they know it was not LLM generated?

The false claim was added 7 Nov 2022 [1] while chatgpt wasn't released until 30 Nov 2022. [1] https://en.wikipedia.org/w/index.php?title=Eugen_Rochko&diff...

Thanks!

Re: Generative AI and Wikipedia editing: What we learned in 2025

#124
post #55

Earlier quoted context omitted.

I’d like to introduce you to the internet. There’s a reason /s was a big thing, one persons obvious sarcasm is (almost tautologically) another persons true statement of opinion.

Thanks. I wasn’t aware of that.

It took me a minute to realise you were joking too! :)

Re: Generative AI and Wikipedia editing: What we learned in 2025

#125
post #109

Earlier quoted context omitted.

Doesnt really matter the reason why does it? If more people use it then it is more popular. Simple metric.

Well, kinda ... Approximately 3.2m Americans drink Hard Kombucha. Approximately 5m Americans use Cocaine. Simple metric.

Right, so cocaine is more popular than Kombucha. Whats your point?

Re: Generative AI and Wikipedia editing: What we learned in 2025

#126
post #125

Earlier quoted context omitted.

Well, kinda ... Approximately 3.2m Americans drink Hard Kombucha. Approximately 5m Americans use Cocaine. Simple metric.

Right, so cocaine is more popular than Kombucha. Whats your point?

{* whoosh * -;-O} - actual statistical output

Re: Generative AI and Wikipedia editing: What we learned in 2025

#127
post #89
post #58

Earlier quoted context omitted.

> but if you dig into the source you realize it's coming from a crank. It is a dark sunday afternoon, Bob Park is sitting on his sofa as usual, drunk as usual, suddenly the TV reveals to him there to be something called the Paranormal (Twilight Zone music) ..instantly Bob knows there are no such things and adds a note to the incomprehensible mess of notes that one day will become his book. He downs one more Budweiser…

Curious what the point you're making here is. I don't know anything at all about Bob Park and whether he is a crank. But if you make your career doing the admirable work of debunking pseudo-science and nonsense theories, you would necessarily be linked to in discussions of those theories very, very frequently. So maybe that's not a good description of him. But the link you posted is hardly dispositive.

The pseudo science of corporate values is a contradiction in terms invented by HR ladies who drink tea for a living. People who believe such things also believe in aliens and have theories about vegetarian tigers.

You are now debunked.

(This comment is intentionally stupid, useless and the author knows nothing about the topic)

Re: Generative AI and Wikipedia editing: What we learned in 2025

#128

So, AI spam can degrade quality. But ... isn't this with regards to Wikipedia a much more general problem? Usually revisions are approved manually by real people. This already can be negative; takes a lot of time; no guarantee that new information is true but old information can be wrong too. To me it seems more as if the problem has much more to do with the quality control problems of wikipedia itself. Yes, AI spam…

"Usually revisions are approved manually by real people."

almost never the case besides a select few articles that get heavily vandalized and require all edits to be approved. otherwise, anyone can edit Wikipedia at any time, which famously is the point of Wikipedia

Re: Generative AI and Wikipedia editing: What we learned in 2025

#129
post #3

So, a small proportion of articles were detected as bot-written, and a large proportion of those failed validation. What if in fact a large proportion of articles were bot-written, but only the unverifiable ones were bad enough to be detected?

Human editors, I suspect, would pick up the "tells" of generated text, although as we know, there's a lot of false positives in that space. But it looks like Pangram is a text classifying NN trained using a technique where they get a human to write a body of text on a subject, and then get various LLMs to write a body of text on the same subject, which strikes me as a good way to approach the problem. Not that I'm in…

If you know what you're looking for -- and no, that does not mean em-dashes, they're a low-signal indicator and people get precious about them -- your false positive rate can get pretty low. See Russell et al, "People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text."

Anecdotally speaking, I'm heavily involved in AI cleanup work on Wikipedia. At this point, I can read article text typical of AI, guess within 6-12 months when that edit was made based on differences in LLM output over the years, and usually be right. (e.g., "delve" was common through early 2024 until GPT-4o killed it; I've seen theories that it's not a GPT-4o thing but rather the result of people avoiding the word when it became a meme, but its frequency dropped off hard even in edits that clearly were not reviewed or edited at all.) Given that Wikipedia has accumulated billions of edits over 25+ years, that says a lot.

Re: Generative AI and Wikipedia editing: What we learned in 2025

#130
post #38

The title I've chosen here is carefully selected to highlight one of the main points. It comes (lightly edited for length) from this paragraph: Far more insidious, however, was something else we discovered: More than two-thirds of these articles failed verification. That means the article contained a plausible-sounding sentence, cited to a real, relevant-sounding source. But when you read the source it’s cited to, th…

People here are claiming that this is true of humans as well. Apart from the fact that bad content can be generated much faster with LLMs, what's your feeling about that criticism? It's there any measure of how many submissions before LLMs make unsubstantiated claims? Thank you for publishing this work. Very useful reminder to verify sources ourselves!

I have indeed seen that with humans as well, including in conference papers and medical journals. The reference citations in papers is seen by many authors as another section they need to fill to get their articles accepted, not as a natural byproduct of writing an article.
Post reply on HN