Live data from Hacker News

Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

m.wikimediafoundation.org

181–190 of 192 posts

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#181

Earlier quoted context omitted.

You make a fair point. I'm not rubbishing Wikipedia, just questioning the supposed USP. I would also point out in response to your argument that a Wikipedia article and a set of search results are apples and oranges. The article is written once then modified or evolved occasionally by (almost exclusively) humans, but very frequently read. It is intended to be intelligible, being structured and based in natural langua…

The other big problem is that curating search results is inherently about prioritising a position rather than establishing a sourced and reasonably neutral version of the truth. I'm imagining the edit wars and debates that take place on contentious wordings or facts in some parts of Wikipedia, but on a much wider scale involving hundreds of SEO consultants each aware that changing a particular criterion will have a q…

Wikipedia already curates links to some extent on every page under "External Links". So there is a seed there.

And even the page text is not immune from the problem you describe. Grading and prioritizing sources is a fundamental part of producing a "reasonably neutral version of the truth." It's what determines what gets cited and how prominently it influences the article.

So while I wouldn't equate text and links in terms of the difficulty of managing POV-neutrality, I would say they sit on a spectrum.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#182
post #47
post #46

Earlier quoted context omitted.

I think facebook should be able to build a search engine. Don't know why they don't have one yet

Facebook wants a single platform to provide the Web through. If they built a search engine, it would only search Facebook.

They already have a FaceBook search engine. I noticed when I start typing a friend's name in the search bar, my friend matches appear below pages that I haven't liked/discussions or if I search an artist I liked, other results appear above it in the quick search drop down. I haven't noticed sponsored search results but I'm sure they're working on it.

Example of a search query URL. Verified pages appear to get more weight. https://www.facebook.com/search/top/?q=president

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#183
post #122

Earlier quoted context omitted.

Do a Google search for anything even slightly obscure, and you're likely to find the first page or so of results filled with highly-SEO'd sites that offer little in the way of deep, detailed content. The smaller sites which do have that content, but just haven't been SEO'd much, have been eclipsed. They're still there, but rendered nearly inaccessible. Interesting discussion on a search engine that does sort of the o…

I hear this a lot, but I'd love to see an example (especially including the sites that should rank).

One example I've experienced is searching for iphone jailbreak related stuff. Perhaps that's to be expected.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#184
post #165
post #147

Earlier quoted context omitted.

This. I don't work in this area, but I'd say that 99% of the time I hear someone complaining about how Google is favoring sites that pay for advertising over them I find that they are making these incredibly basic errors. For me, http://www.linguaquote.com/ gives a 404. It's only when I go to https://www.linguaquote.com/ that it works.

That's all well and good, and thanks for checking, but I wasn't referring to that site in the parent comment. The other site does have SSL enabled, but only recently and via Let's Encrypt. The issue is much longer standing than this. So I'm afraid it's not quite as simple as you make out in this case. In other news, I've just pinpointed the missing line in the recently changed nginx config for Linguaquote; the http:/…

Is it bad form to quote myself?

I'd love to see an example (especially including the sites that should rank).

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#186
As a former NLP engineer and former wikiHow engineer, I have some perspective on this. Google has included more and more information from Wikipedia. Furthermore, Google includes snippets of external websites in the knowledge box on more and more pages.

How long will be until can algorithmically generate its own Wikipedia articles? Wikipedia relies upon coming to its site for contributions and donations. Without search, Wikipedia risks being subsumed by Google. They have a difficult position of thinking about the future without pissing off Google.

Computers are getting more and more powerful. Wikipedia needs to do stay relevant. I think this is the right decision.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#187
post #152
post #124

Earlier quoted context omitted.

Why do they need 4 team leads for 8 engineers?

WMF projects are usually open to volunteer help

So? I've found 1:10 to be a good ratio between leads and grunts. 1:5 if lead is technical and also contributes work and not just leadership/management. And that's when those 10 (or 5) people are working full time.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#188
post #184
post #165

Earlier quoted context omitted.

That's all well and good, and thanks for checking, but I wasn't referring to that site in the parent comment. The other site does have SSL enabled, but only recently and via Let's Encrypt. The issue is much longer standing than this. So I'm afraid it's not quite as simple as you make out in this case. In other news, I've just pinpointed the missing line in the recently changed nginx config for Linguaquote; the http:/…

Is it bad form to quote myself? I'd love to see an example (especially including the sites that should rank).

I put the link to the freelance site in my profile, which I only mentioned in my reply to cpeterso above - apologies for not making it clearer.

Will leave it there a bit longer in case you do come back to this thread as still genuinely interested in your opinion on the matter.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#189

If nothing else, I hope it improves the currently abysmal search features for Wikipedia today.

turns out that is the only thing this project is about. There is no web crawler, there is no external content. The grant and the money WMF is spending are being spent to improve internal search at wikipedia.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#190
post #110

As a former Wikia employee, I am somewhat of a MediaWiki insider. I sped Wikia's search engine up by several orders of magnitude and then went on to pilot a number of NLP/machine learning initiatives in the company. Jimmy Wales' already tried to make a "Google Killer" ten years ago. It was tilting at windmills to say the least. Letting individuals help manage algorithmic search results was harder than you could imagi…

How did you arrive at 20 million? This sounds like one of those "technically true" facts that are cooked up for investors. http://wikis.wikia.com/wiki/List_of_Wikia_wikis puts the combined total of the top 1,000 wikis (in all languages) at 12.4m.

20 million pages, not wikis -- sorry if I mistyped?
Post reply on HN