Live data from Hacker News

Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

m.wikimediafoundation.org

151–160 of 192 posts

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#151
Wikipedia main asset seems based around user contributions and human interactions not programming or hard algorithms, this seems quite a leap into another field with not much money.

Bing cost MS $5.5 Billion in their field of expertise

http://www.geek.com/news/bing-has-cost-microsoft-5-5-billion...

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#152
post #124
post #112

Earlier quoted context omitted.

They plan to hire 8 engineers, 2 data analysts, and 4 team leads. Relevant details on page 8. https://twitter.com/dtunkelang/status/699049083806167040

Why do they need 4 team leads for 8 engineers?

WMF projects are usually open to volunteer help

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#154
This is like competing with Intel in CPU server market, where he has 98% market share.

So they are trying to compete using 2.5 mil dollars with software backed by multi billion dollars, hundreds thousands of servers, tons of data, thousands of developers, ML integration etc.?

Good luck with that. Many tried backed by x times more resources than this 2.5 mil, unfortunately all failed.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#155
post #66

This was/is actually an extremely controversial project. The corporation (basically the Executive Director) pursued the grant and the idea without soliciting input or really disclosing it to the community of editors, and eventually one of the community-elected trustees was removed for questioning the lack of transparency. The community has a long list of software improvements that they'd like to see to the core platf…

> and eventually one of the community-elected trustees was removed for questioning the lack of transparency This was a hypothesis about his removal when he was initially was removed, and has since been refuted by multiple sources.

[citation needed]

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#156
post #66

This was/is actually an extremely controversial project. The corporation (basically the Executive Director) pursued the grant and the idea without soliciting input or really disclosing it to the community of editors, and eventually one of the community-elected trustees was removed for questioning the lack of transparency. The community has a long list of software improvements that they'd like to see to the core platf…

Wikipedia needs to solve it's management issues before it starts developing new software...

I've vouched for this comment.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#157
Computing is so cheap now, google isn't going to be so dominant on text search for long. Their money is needed for video and pictures and audio, but the text internet can be cached whole by small entities now.

Maybe Wikipedia should launch a video encyclopedia to try to provide a 5 minute video of every article, for people who like videos more than reading.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#158
post #95

Earlier quoted context omitted.

> Curation doesn't scale and often assumes a single unifying ontology Wikipedia is a pretty big exception to that assertion. Perhaps DMOZ (a clone of Yahoo circa 1996) is not the only way to do curation. Perhaps Wikipedia could apply what has worked for Wikipedia, i.e. develop a set of POV-neutral criteria for organizing collections of links and then invite everyone to participate. It's really easy to be negative. Bu…

You make a fair point. I'm not rubbishing Wikipedia, just questioning the supposed USP. I would also point out in response to your argument that a Wikipedia article and a set of search results are apples and oranges. The article is written once then modified or evolved occasionally by (almost exclusively) humans, but very frequently read. It is intended to be intelligible, being structured and based in natural langua…

The other big problem is that curating search results is inherently about prioritising a position rather than establishing a sourced and reasonably neutral version of the truth.

I'm imagining the edit wars and debates that take place on contentious wordings or facts in some parts of Wikipedia, but on a much wider scale involving hundreds of SEO consultants each aware that changing a particular criterion will have a quantifiable impact on their clients' bottom line. It doesn't sound like it would be fun to police.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#159

I've read a log of the inside wikimedia links here, and I'm confused about all the talk of gnashing of teeth and rending of cloth. This is controversial because some want to pay down technical debt rather than have a small team do knowledge graph search? Okay...

It's not a small team compared to the size if the organisation. And you mischaracterise the situation: this is a problem with engagement, transparency and openness.
Post reply on HN