Live data from Hacker News

In Sweden, Sverker Johansson and His 'Bot' Have Created 2.7M Wikipedia Articles

online.wsj.com

51–60 of 78 posts

Re: In Sweden, Sverker Johansson and His 'Bot' Have Created 2.7M Wikipedia Articles

#51
post #29
post #3

I'm not very in-tune with Wikipedia's culture (which, I've read, is very nuanced and rigid[1]), but I really don't see why this is a bad thing, given the information in the articles is accurate (and the article gives the impression that glitches are rare). If nobody else was going to create an article about some species of butterfly, I don't see why adding that information would be harmful to Wikipedia. Does it make…

At least part of the problem is that he's generating what one might call 'info trash': he's taking highly structured information from databases, and turning it into natural-language prose, a data source of less value since it's less structured. These prose versions are now going to steadily fall out of sync with the original databases, be much more prominent in Wikipedia and Google, diverge from each other, be harder…

But now you are arguing the "it would be better if you did Y instead of X, so therefore stop doing X" fallacy. It's not like he is pillaging the original databases and leaves them burning. They are still there and his copying of data from them doesn't make them worse. You or anyone else are welcome to spend your free time exporting and merging the databases into Freebase.

Re: In Sweden, Sverker Johansson and His 'Bot' Have Created 2.7M Wikipedia Articles

#52

I think this is going to be more common in the future. Not only for Wikipedia, but for every page that serves some kind of content. Perhaps even sites like 9gag or similar could start out with some computer generated memes ;)

Ha! yeah I thought about a hack that just gamified random scraped words with memes and see what rises to the top and receptive to humans.

But still its two different horses, Wikipedia is suppose to be a central repository of community created knowledge. The veracity of the information and expertise of authors is what secured it as a credible popular source.

Would we say the same thing if this was a central repository of just Stubs that were computer generated from the get-go?

Re: In Sweden, Sverker Johansson and His 'Bot' Have Created 2.7M Wikipedia Articles

#53
post #28

The only thing that pisses me off is that when I contribute I'm expected to write half a site in addition to such an infobox, otherwise a mod comes along and movees it into my sandbox. That's the main reason I always stopped editing Wikipedia right away after I tryed.

Exactly, agreed this is why fringe projects are created for more opinionated subjective matter or cultural matters.

If you love contributing to wikipedia style knowledge, @Localwiki is often the anti-wikipedia, about local relevant knowledge but colloquial prose is usually OK as long as its factual and not malicious.

http://localwiki.org/

Increasing the barrier to entry, generally never helps inclusion or is newbie friendly.

Re: In Sweden, Sverker Johansson and His 'Bot' Have Created 2.7M Wikipedia Articles

#54
post #36

Why is he doing this in Swedish and two versions of Filipino? He apparently speaks English, so I assume he could adapt his code to do English entries as well.

I guess he only has a finite amount of time and energy. Converting the bot might be the easiest part, but then, you also need to convince the English 'pedians.

Also I would imagine many of the articles in English may have been already seeded as in general Swedish and Tagalog are more esoteric.

Re: In Sweden, Sverker Johansson and His 'Bot' Have Created 2.7M Wikipedia Articles

#55
post #29
post #3

I'm not very in-tune with Wikipedia's culture (which, I've read, is very nuanced and rigid[1]), but I really don't see why this is a bad thing, given the information in the articles is accurate (and the article gives the impression that glitches are rare). If nobody else was going to create an article about some species of butterfly, I don't see why adding that information would be harmful to Wikipedia. Does it make…

At least part of the problem is that he's generating what one might call 'info trash': he's taking highly structured information from databases, and turning it into natural-language prose, a data source of less value since it's less structured. These prose versions are now going to steadily fall out of sync with the original databases, be much more prominent in Wikipedia and Google, diverge from each other, be harder…

Some people (a fairly significant number) find it much easier to parse information presented in English sentences as opposed to the presentation forms typical in structured data (often table form, with some kind of filtering).

In many cases, it is very rare for facts like those presented in botanical databases to change: it often means a plant has been recategorized, which is a non-trivial thing to do. It is entirely appropriate for this to be handled manually, given how rare it is.

Your arguments about it being better to work on a structured version present a false dichotomy: it isn't Wikipedia OR a structured version, it is Wikipedia AND structured data that need to be improved.

Re: In Sweden, Sverker Johansson and His 'Bot' Have Created 2.7M Wikipedia Articles

#57
post #3

I'm not very in-tune with Wikipedia's culture (which, I've read, is very nuanced and rigid[1]), but I really don't see why this is a bad thing, given the information in the articles is accurate (and the article gives the impression that glitches are rare). If nobody else was going to create an article about some species of butterfly, I don't see why adding that information would be harmful to Wikipedia. Does it make…

Looking at [1], the following can be found listed under sources:

^ , 1901 In: Bull. Div. Agrostol. U.S.D.A. 24: 26

Maybe he should have done more testing before spamming Wikipedia with mess like that.

[1] https://sv.wikipedia.org/wiki/Eutriana_repens

Re: In Sweden, Sverker Johansson and His 'Bot' Have Created 2.7M Wikipedia Articles

#58
post #7

Why doesn't this guy just package up the database, and autogenerate the page if the page is blank. That would have more utility than mixing human and bot generated articles.

Unless there's some nuance that I'm missing, it sounds like that that is exactly what he's doing.

I don't think that's what grizzles is saying, I think the idea was to patch the wiki software to generate these article stubs on the fly from the actual source instead of batch importing them once.

Sounds like a good idea to me, but I don't know the Wikipedia culture so they might have reasons against that.

Re: In Sweden, Sverker Johansson and His 'Bot' Have Created 2.7M Wikipedia Articles

#59
post #39
post #35

Earlier quoted context omitted.

> If Wikipedia were the only website I'd agree, but I don't find it useful for Wikipedia to have stubs that are simple copies of other freely available sources (especially more authoritative ones), without some kind of synthesis or value-add. The value-add is that I know about Wikipedia, but not about whatever more authoritative botanical site you're mentioning.

That is why we search. With the nearly empty wikipedia page in place, you might never be driven to find the better source.

If I'm not driven to find a better source, then presumably I was satisfied with the information on Wikipedia. Quick reference for the most relevant information is what Wikipedia is good at.

Re: In Sweden, Sverker Johansson and His 'Bot' Have Created 2.7M Wikipedia Articles

#60
post #8
post #3

I'm not very in-tune with Wikipedia's culture (which, I've read, is very nuanced and rigid[1]), but I really don't see why this is a bad thing, given the information in the articles is accurate (and the article gives the impression that glitches are rare). If nobody else was going to create an article about some species of butterfly, I don't see why adding that information would be harmful to Wikipedia. Does it make…

I think the problem is that the entries are very low-quality and are being produced in such great quantities that it will be hard for anyone to turn them into passable articles. The purpose of Wikipedia is not to be a collection of all the factual information it gather. You or I might wish for it to be one, but that isn't what its creators mean for it to be. Each individual article is expected to meet Wikipedia's gui…

I think these stubs have the benefit of reducing the barrier to entry for future contributors. If I search for something and find a stub, I can easily throw in even one sentence, and the article is incrementally improved. Whereas if there's no article, I am much less inclined to make a new one.
Post reply on HN