Live data from Hacker News

Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

m.wikimediafoundation.org

101–110 of 192 posts

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#101
post #95

Earlier quoted context omitted.

1) Public curation mechanisms for quality; The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc. 2) Transparency, telling users exactly how the inf…

> Curation doesn't scale and often assumes a single unifying ontology Wikipedia is a pretty big exception to that assertion. Perhaps DMOZ (a clone of Yahoo circa 1996) is not the only way to do curation. Perhaps Wikipedia could apply what has worked for Wikipedia, i.e. develop a set of POV-neutral criteria for organizing collections of links and then invite everyone to participate. It's really easy to be negative. Bu…

You make a fair point. I'm not rubbishing Wikipedia, just questioning the supposed USP. I would also point out in response to your argument that a Wikipedia article and a set of search results are apples and oranges.

The article is written once then modified or evolved occasionally by (almost exclusively) humans, but very frequently read. It is intended to be intelligible, being structured and based in natural language. It has a very well defined scope within a flat namespace, and often clear relations to multiple formal ontologies. It is structured to be consumed in part or in whole, and may contain rich media and strong supporting contextual information (related pages).

By contrast a search result summarizes a set of potential information sources that may answer a search query in whole or in part, to various definitions of "answer". It is generally written once, by a computer, and thrown away after some period of caching. It is intended to be concise. Each component result has relatively poor context, relying upon the searcher to interpret timeliness, authority, notability, uniqueness, comprehensibility, etc. with the limited information presented, typically a very short content excerpt. It is structured to be scanned, classically in a ranked fashion from "best hit" to "worst hit", and is generally a wall of text.

Wikipedia successfully attracts people to contribute to the former, but the latter - where the information product is generated on the fly and lasting impact is amorphous (nothing particularly concrete for contributors to point to and say "I did that! Warm and fuzzies!") - is a very different beast.

I too believe there is room for innovation ... there are potentially low hanging fruit like inter-linguistic semantic queries (not keyword search) ... but there are no such key problem areas identified in the paper's summary.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#102
post #51

Earlier quoted context omitted.

Care to explain? Do you have some links/sources?

Sure. I'll start off with the following email from Liam Wyatt: https://lists.wikimedia.org/pipermail/wikimedia-l/2016-Febru... The grant application you are looking at was only revealed due to a MASSIVE amount of controversy and pressure within the Wikimedia Foundation. The community representative (James Heilman) on the board was let go the other day, in part because of concerns around this grant. You might want to…

It may be of little importance vs your excellent references and what they show but... does anyone else notice how she worded that message is just... so... weird? The wording comes off like a combination of academia, PR, and email scams to me. Just straight BS that no normal, caring person in a mission-oriented organization should ever say.

I mean, there's certainly styles I'm unfamiliar with. I'm always open to new experiences. Could be the case here. Hers just instantly set off red flags in my intuition. I hope she didn't always write like that as it might mean whoever brought her in either fell for a con or were part of it.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#103
post #38

Summary of the approach (p10): "1) Public curation mechanisms for quality; 2) Transparency, telling users exactly how the information originated; 3) Open data access to metadata, giving users the exact date source of the information; 4) Protected user privacy, with their searching protected by strict privacy controls; 5) No advertising, which assures the free flow of information and a complete separation from commerc…

1) Public curation mechanisms for quality; The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc. 2) Transparency, telling users exactly how the inf…

> We have duckduckgo already

DuckDuckGo a meta-search-engine! It relies mainly on Yahoo Boss API which uses Bing search (for most countries)! Yahoo Boss API turned from free to expensive in early 2015 and the future of Yahoo (tech company, not Alibaba stock) is uncertain.

We definitely need more search engines, only 6-7 exist that cover a wider range (international). Search on HN to retrieve the list, we had this discussion before.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#104
post #38

Summary of the approach (p10): "1) Public curation mechanisms for quality; 2) Transparency, telling users exactly how the information originated; 3) Open data access to metadata, giving users the exact date source of the information; 4) Protected user privacy, with their searching protected by strict privacy controls; 5) No advertising, which assures the free flow of information and a complete separation from commerc…

"Public Curation" doesn't make quality. It makes a mob-rule system where only the most popular ideas flourish.

Nonsense. It means content will be filtered through the lens of one or more individuals. The results vary dramatically. Mob-rule is one possibility. Yahoo Directory used to be a great example of mid-level of quality where it gave nice start and obscure stuff people overlooked. On high end, the link below shows Stanford Encyclopedia of Philosophy set a pretty awesome precedent for high-quality curation:

https://news.ycombinator.com/item?id=10266103

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#106
post #103

Earlier quoted context omitted.

1) Public curation mechanisms for quality; The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc. 2) Transparency, telling users exactly how the inf…

> We have duckduckgo already DuckDuckGo a meta-search-engine! It relies mainly on Yahoo Boss API which uses Bing search (for most countries)! Yahoo Boss API turned from free to expensive in early 2015 and the future of Yahoo (tech company, not Alibaba stock) is uncertain. We definitely need more search engines, only 6-7 exist that cover a wider range (international). Search on HN to retrieve the list, we had this dis…

I agree with you, but this quote was in the context of the claim of a privacy USP. You have taken it out of context.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#107
post #66

This was/is actually an extremely controversial project. The corporation (basically the Executive Director) pursued the grant and the idea without soliciting input or really disclosing it to the community of editors, and eventually one of the community-elected trustees was removed for questioning the lack of transparency. The community has a long list of software improvements that they'd like to see to the core platf…

[deleted]

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#108
post #97

This is a very, very, very small amount of money if you want to build a search engine, let alone one "to rival Google" (source?). Looks like the goals are realistic, though - look how wikipedia search could be extended beyond results from wikipedia.org, build some test sets. And get a better idea what it really is that is supposed to be built.

Cuil spent about $30M before they went bust. By the time they went under, due to a total lack of a revenue model, they had a halfway decent search engine that did its own web crawl. So that's a data point on how much it costs.

If only Cuil came about after the Snowden revelations, they might have been able to make a real name for themselves.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#109
post #105

That's how much 5 qualified software engineers would cost to employ for a year (gross, including compensation, benefits, payroll taxes, office space, hardware, etc, and that's on the low end of the range). Good luck with that.

It costs half a million to employ one software engineer for a year?

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#110

As a former Wikia employee, I am somewhat of a MediaWiki insider. I sped Wikia's search engine up by several orders of magnitude and then went on to pilot a number of NLP/machine learning initiatives in the company. Jimmy Wales' already tried to make a "Google Killer" ten years ago. It was tilting at windmills to say the least. Letting individuals help manage algorithmic search results was harder than you could imagi…

How did you arrive at 20 million? This sounds like one of those "technically true" facts that are cooked up for investors. http://wikis.wikia.com/wiki/List_of_Wikia_wikis puts the combined total of the top 1,000 wikis (in all languages) at 12.4m.
Post reply on HN