Live data from Hacker News

Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

m.wikimediafoundation.org

111–120 of 192 posts

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#111
post #97

Earlier quoted context omitted.

Cuil spent about $30M before they went bust. By the time they went under, due to a total lack of a revenue model, they had a halfway decent search engine that did its own web crawl. So that's a data point on how much it costs.

If only Cuil came about after the Snowden revelations, they might have been able to make a real name for themselves.

That's been a real boost to DuckDuckGo.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#112
post #105

That's how much 5 qualified software engineers would cost to employ for a year (gross, including compensation, benefits, payroll taxes, office space, hardware, etc, and that's on the low end of the range). Good luck with that.

They plan to hire 8 engineers, 2 data analysts, and 4 team leads. Relevant details on page 8.

https://twitter.com/dtunkelang/status/699049083806167040

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#113

most of my google searches include a wikipedia result on the first page. I would estimate this could reduce Google's web search revenue by upwards of 40% worldwide.

Google doesn't make that much money on research quieries. It's mostly your other queries are what enables them to sell ads.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#114
post #109
post #105

That's how much 5 qualified software engineers would cost to employ for a year (gross, including compensation, benefits, payroll taxes, office space, hardware, etc, and that's on the low end of the range). Good luck with that.

It costs half a million to employ one software engineer for a year?

Cost per employee is much greater than just their salary.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#116
post #103

Earlier quoted context omitted.

1) Public curation mechanisms for quality; The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc. 2) Transparency, telling users exactly how the inf…

> We have duckduckgo already DuckDuckGo a meta-search-engine! It relies mainly on Yahoo Boss API which uses Bing search (for most countries)! Yahoo Boss API turned from free to expensive in early 2015 and the future of Yahoo (tech company, not Alibaba stock) is uncertain. We definitely need more search engines, only 6-7 exist that cover a wider range (international). Search on HN to retrieve the list, we had this dis…

BOSS is being discontinued on March 31st, so I imagine that DDG has some sort of idea what they're going to be doing after that date.

Any insight?

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#117

Earlier quoted context omitted.

I'm pretty sure that the spammer argument is just an excuse used by Google to allow them to keep their business practices out of public scrutiny. Google search results are biased in favour of content produced by those who have money and power. Google ranks everything based on popularity - Not based on quality. Popularity and quality are two independent concepts and not necessarily related. That's something which Wiki…

I'd ask you to cite your claims, but we both know you can't. It's a pity your issues with Google cause you to pollute discussions with BS.

The primary measurement used for Google's PageRank algorithm is the number and "value" of backlinks that a page has. The "value" of a backlink is determined by the cumulative number of descendent sub-backlinks that it has. This is common knowledge among SEO professionals.

Basically, when judging quality, Google is making assumptions like: "This page has a lot of backlinks, and those backlinks themselves have a lot of backlinks... Therefore this page is of high quality." This approach puts all the power in the hands of content providers (bloggers) who are funded by big companies (or well-funded startups) and who serve the interests of those companies.

Google wrongly assumes that content-providers serve the interest of consumers and that they can be trusted - Which is not the case.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#118
post #66

This was/is actually an extremely controversial project. The corporation (basically the Executive Director) pursued the grant and the idea without soliciting input or really disclosing it to the community of editors, and eventually one of the community-elected trustees was removed for questioning the lack of transparency. The community has a long list of software improvements that they'd like to see to the core platf…

Wikipedia needs to solve it's management issues before it starts developing new software...

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#119

Earlier quoted context omitted.

Sure. I'll start off with the following email from Liam Wyatt: https://lists.wikimedia.org/pipermail/wikimedia-l/2016-Febru... The grant application you are looking at was only revealed due to a MASSIVE amount of controversy and pressure within the Wikimedia Foundation. The community representative (James Heilman) on the board was let go the other day, in part because of concerns around this grant. You might want to…

It may be of little importance vs your excellent references and what they show but... does anyone else notice how she worded that message is just... so... weird? The wording comes off like a combination of academia, PR, and email scams to me. Just straight BS that no normal, caring person in a mission-oriented organization should ever say. I mean, there's certainly styles I'm unfamiliar with. I'm always open to new e…

[deleted]

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#120
post #103

Earlier quoted context omitted.

> We have duckduckgo already DuckDuckGo a meta-search-engine! It relies mainly on Yahoo Boss API which uses Bing search (for most countries)! Yahoo Boss API turned from free to expensive in early 2015 and the future of Yahoo (tech company, not Alibaba stock) is uncertain. We definitely need more search engines, only 6-7 exist that cover a wider range (international). Search on HN to retrieve the list, we had this dis…

BOSS is being discontinued on March 31st, so I imagine that DDG has some sort of idea what they're going to be doing after that date. Any insight?

They are using Yandex now with a bit of their own crawling.
Post reply on HN