Live data from Hacker News

Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

m.wikimediafoundation.org

71–80 of 192 posts

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#71
I'm not kidding when I say that if they want to know where to spend the $2.5m I would start with cleaning up their core codebase. IMO Mediawiki open source code is a disaster.

EDIT: Not because it's written in PHP. Because it's architected poorly.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#72
post #66

This was/is actually an extremely controversial project. The corporation (basically the Executive Director) pursued the grant and the idea without soliciting input or really disclosing it to the community of editors, and eventually one of the community-elected trustees was removed for questioning the lack of transparency. The community has a long list of software improvements that they'd like to see to the core platf…

A bit more detail is shared in this comment below: https://news.ycombinator.com/user?id=chris_wot

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#73
post #72
post #66

This was/is actually an extremely controversial project. The corporation (basically the Executive Director) pursued the grant and the idea without soliciting input or really disclosing it to the community of editors, and eventually one of the community-elected trustees was removed for questioning the lack of transparency. The community has a long list of software improvements that they'd like to see to the core platf…

A bit more detail is shared in this comment below: https://news.ycombinator.com/user?id=chris_wot

More specific link: https://news.ycombinator.com/item?id=11101262

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#74
post #72

Earlier quoted context omitted.

A bit more detail is shared in this comment below: https://news.ycombinator.com/user?id=chris_wot

More specific link: https://news.ycombinator.com/item?id=11101262

Better off here:

https://news.ycombinator.com/item?id=11101164

:-)

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#75
post #38

Summary of the approach (p10): "1) Public curation mechanisms for quality; 2) Transparency, telling users exactly how the information originated; 3) Open data access to metadata, giving users the exact date source of the information; 4) Protected user privacy, with their searching protected by strict privacy controls; 5) No advertising, which assures the free flow of information and a complete separation from commerc…

1) Public curation mechanisms for quality;

The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc.

2) Transparency, telling users exactly how the information originated;

Most search engines include a URL, I can see a [crawldate] button like the [cache] or [translate] buttons on each hit adding some information, but this will be of dubious additional utility for most searches.

3) Open data access to metadata, giving users the exact date source of the information;

As above.

4) Protected user privacy, with their searching protected by strict privacy controls;

We have duckduckgo already, friends are welcome but it's hardly a unique offering nor a trustworthy one given Snowden's revelations regarding the scale of systematic 5 eyes traffic monitoring/recording.

5) No advertising, which assures the free flow of information and a complete separation from commercial interests;

DDG or Google or Bing with plugins can supply this. Not ground breaking.

6) Internalization, which emphasizes community building and the sharing of information instead of a top-down approach.

This is so amorphous as to be a non-point.

So out of six points, 2 things (33%) are only useful in edge cases, 1 thing (16%) is too vague to be useful, and the other 3 things (50%) are currently implemented by others and have been tried before.

I would like to see the input of the former Blekko guys on this, https://news.ycombinator.com/user?id=ChuckMcM + https://news.ycombinator.com/user?id=greglindahl

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#76

I'm not kidding when I say that if they want to know where to spend the $2.5m I would start with cleaning up their core codebase. IMO Mediawiki open source code is a disaster. EDIT: Not because it's written in PHP. Because it's architected poorly.

It's funny you should mention that. That was a point that apparently a number of WMF staff expressed, and it was apparently ignored.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#77
post #59
post #47

Earlier quoted context omitted.

Facebook wants a single platform to provide the Web through. If they built a search engine, it would only search Facebook.

Facebook search is broken, twitter search works much better.

It's intentionally broken.

Despite many real (though also some exaggerated) counter-examples, Facebook does have features to protect privacy.

One of those things is you generally can't get very useful results from search unless you're friends with someone, or a friend of a friend (depending on the user's privacy settings). You can't see general trends, or even search for every person with a given first and last name in an area, for example.

It used to be more open, but they've heavily restricted the breadth of data returned from searches within the past few years.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#78
post #38

Summary of the approach (p10): "1) Public curation mechanisms for quality; 2) Transparency, telling users exactly how the information originated; 3) Open data access to metadata, giving users the exact date source of the information; 4) Protected user privacy, with their searching protected by strict privacy controls; 5) No advertising, which assures the free flow of information and a complete separation from commerc…

1) Public curation mechanisms for quality; The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc. 2) Transparency, telling users exactly how the inf…

Duckduckgo is ad free? I never knew this. How do they make money?

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#79

Earlier quoted context omitted.

I'm pretty sure that the spammer argument is just an excuse used by Google to allow them to keep their business practices out of public scrutiny. Google search results are biased in favour of content produced by those who have money and power. Google ranks everything based on popularity - Not based on quality. Popularity and quality are two independent concepts and not necessarily related. That's something which Wiki…

I'd ask you to cite your claims, but we both know you can't. It's a pity your issues with Google cause you to pollute discussions with BS.

Do a Google search for anything even slightly obscure, and you're likely to find the first page or so of results filled with highly-SEO'd sites that offer little in the way of deep, detailed content. The smaller sites which do have that content, but just haven't been SEO'd much, have been eclipsed. They're still there, but rendered nearly inaccessible.

Interesting discussion on a search engine that does sort of the opposite of Google: https://news.ycombinator.com/item?id=3910304

Post reply on HN