EDIT: Not because it's written in PHP. Because it's architected poorly.
Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
71–80 of 192 posts
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#72This was/is actually an extremely controversial project. The corporation (basically the Executive Director) pursued the grant and the idea without soliciting input or really disclosing it to the community of editors, and eventually one of the community-elected trustees was removed for questioning the lack of transparency. The community has a long list of software improvements that they'd like to see to the core platf…
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#73This was/is actually an extremely controversial project. The corporation (basically the Executive Director) pursued the grant and the idea without soliciting input or really disclosing it to the community of editors, and eventually one of the community-elected trustees was removed for questioning the lack of transparency. The community has a long list of software improvements that they'd like to see to the core platf…
A bit more detail is shared in this comment below: https://news.ycombinator.com/user?id=chris_wot
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#74Earlier quoted context omitted.
A bit more detail is shared in this comment below: https://news.ycombinator.com/user?id=chris_wot
More specific link: https://news.ycombinator.com/item?id=11101262
https://news.ycombinator.com/item?id=11101164
:-)
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#75Summary of the approach (p10): "1) Public curation mechanisms for quality; 2) Transparency, telling users exactly how the information originated; 3) Open data access to metadata, giving users the exact date source of the information; 4) Protected user privacy, with their searching protected by strict privacy controls; 5) No advertising, which assures the free flow of information and a complete separation from commerc…
The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc.
2) Transparency, telling users exactly how the information originated;
Most search engines include a URL, I can see a [crawldate] button like the [cache] or [translate] buttons on each hit adding some information, but this will be of dubious additional utility for most searches.
3) Open data access to metadata, giving users the exact date source of the information;
As above.
4) Protected user privacy, with their searching protected by strict privacy controls;
We have duckduckgo already, friends are welcome but it's hardly a unique offering nor a trustworthy one given Snowden's revelations regarding the scale of systematic 5 eyes traffic monitoring/recording.
5) No advertising, which assures the free flow of information and a complete separation from commercial interests;
DDG or Google or Bing with plugins can supply this. Not ground breaking.
6) Internalization, which emphasizes community building and the sharing of information instead of a top-down approach.
This is so amorphous as to be a non-point.
So out of six points, 2 things (33%) are only useful in edge cases, 1 thing (16%) is too vague to be useful, and the other 3 things (50%) are currently implemented by others and have been tried before.
I would like to see the input of the former Blekko guys on this, https://news.ycombinator.com/user?id=ChuckMcM + https://news.ycombinator.com/user?id=greglindahl
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#76I'm not kidding when I say that if they want to know where to spend the $2.5m I would start with cleaning up their core codebase. IMO Mediawiki open source code is a disaster. EDIT: Not because it's written in PHP. Because it's architected poorly.
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#77Earlier quoted context omitted.
Facebook wants a single platform to provide the Web through. If they built a search engine, it would only search Facebook.
Facebook search is broken, twitter search works much better.
Despite many real (though also some exaggerated) counter-examples, Facebook does have features to protect privacy.
One of those things is you generally can't get very useful results from search unless you're friends with someone, or a friend of a friend (depending on the user's privacy settings). You can't see general trends, or even search for every person with a given first and last name in an area, for example.
It used to be more open, but they've heavily restricted the breadth of data returned from searches within the past few years.
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#78Summary of the approach (p10): "1) Public curation mechanisms for quality; 2) Transparency, telling users exactly how the information originated; 3) Open data access to metadata, giving users the exact date source of the information; 4) Protected user privacy, with their searching protected by strict privacy controls; 5) No advertising, which assures the free flow of information and a complete separation from commerc…
1) Public curation mechanisms for quality; The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc. 2) Transparency, telling users exactly how the inf…
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#79Earlier quoted context omitted.
I'm pretty sure that the spammer argument is just an excuse used by Google to allow them to keep their business practices out of public scrutiny. Google search results are biased in favour of content produced by those who have money and power. Google ranks everything based on popularity - Not based on quality. Popularity and quality are two independent concepts and not necessarily related. That's something which Wiki…
I'd ask you to cite your claims, but we both know you can't. It's a pity your issues with Google cause you to pollute discussions with BS.
Interesting discussion on a search engine that does sort of the opposite of Google: https://news.ycombinator.com/item?id=3910304