Live data from Hacker News

Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

m.wikimediafoundation.org

121–130 of 192 posts

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#121
post #38

Summary of the approach (p10): "1) Public curation mechanisms for quality; 2) Transparency, telling users exactly how the information originated; 3) Open data access to metadata, giving users the exact date source of the information; 4) Protected user privacy, with their searching protected by strict privacy controls; 5) No advertising, which assures the free flow of information and a complete separation from commerc…

I'm pretty sure that the spammer argument is just an excuse used by Google to allow them to keep their business practices out of public scrutiny. Google search results are biased in favour of content produced by those who have money and power. Google ranks everything based on popularity - Not based on quality. Popularity and quality are two independent concepts and not necessarily related. That's something which Wiki…

If you think "spam" isn't the defining problem of web search then you've never tried building a search engine. It's 90% of the problem.

Google does take plenty of quality features into account[1]. PageRank is one of course, but that isn't some corporate conspiracy, it's that it is a good feature.

[1] https://moz.com/search-ranking-factors/correlations

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#122

Earlier quoted context omitted.

I'd ask you to cite your claims, but we both know you can't. It's a pity your issues with Google cause you to pollute discussions with BS.

Do a Google search for anything even slightly obscure, and you're likely to find the first page or so of results filled with highly-SEO'd sites that offer little in the way of deep, detailed content. The smaller sites which do have that content, but just haven't been SEO'd much, have been eclipsed. They're still there, but rendered nearly inaccessible. Interesting discussion on a search engine that does sort of the o…

I hear this a lot, but I'd love to see an example (especially including the sites that should rank).

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#123
post #109
post #105

That's how much 5 qualified software engineers would cost to employ for a year (gross, including compensation, benefits, payroll taxes, office space, hardware, etc, and that's on the low end of the range). Good luck with that.

It costs half a million to employ one software engineer for a year?

A software engineer qualified to work on this kind of thing is worth about $350K in combined compensation on the market right now. Typically half of that is base, while the other half is stock and other taxable benefits. The number can be higher. This is the cost of just compensation to the company, excluding the payroll tax. You can, of course, find someone a lot cheaper, but then you'd be a fool to expect the result to be anywhere near as good as what Google can pull off, because if that someone could do what Google can, why would she work for half the compensation instead of applying to Google or FB or whoever pays competitive salaries these days.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#124
post #112
post #105

That's how much 5 qualified software engineers would cost to employ for a year (gross, including compensation, benefits, payroll taxes, office space, hardware, etc, and that's on the low end of the range). Good luck with that.

They plan to hire 8 engineers, 2 data analysts, and 4 team leads. Relevant details on page 8. https://twitter.com/dtunkelang/status/699049083806167040

Why do they need 4 team leads for 8 engineers?

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#127
I've read a log of the inside wikimedia links here, and I'm confused about all the talk of gnashing of teeth and rending of cloth. This is controversial because some want to pay down technical debt rather than have a small team do knowledge graph search?

Okay...

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#128
post #125

Weird; they have anual drives to raise money to keep the site running; would not expect that the'd have 2.5m lying around to do pet projects like this...

Most non-profits, especially the ones that are always asking, usually have a lot of funds.

Before I give, I go to guidestar, hit free preview(they try to trick you into a paying membership), download last few years of 1040's, and see if everything looks copacetic. I look at who is making the most money. Their is usually one person making a very good living. California non-profits are much easier to scrutinize than Deleware non-profits.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#129
post #66

This was/is actually an extremely controversial project. The corporation (basically the Executive Director) pursued the grant and the idea without soliciting input or really disclosing it to the community of editors, and eventually one of the community-elected trustees was removed for questioning the lack of transparency. The community has a long list of software improvements that they'd like to see to the core platf…

There's a new Signpost edition, covering this grant document and also three other internal documents that were leaked.

https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/S...

Post reply on HN