Live data from Hacker News

Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

m.wikimediafoundation.org

91–100 of 192 posts

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#91

Earlier quoted context omitted.

It's funny you should mention that. That was a point that apparently a number of WMF staff expressed, and it was apparently ignored.

There used to be a team dedicated to making MW Core better / cleaner, but that was lost in a re-org earlier in 2015.

That's a darned shame. I'm aware that there are a lot of areas that people want to fix on MW Core.

I get really concerned when I hear that the person who holds the vision and direction for the Wikimedia Foundation didn't really participate in it beforehand, and I get even more concerned when I see that she branches off into proposals for search technology that appear to be far outside the scope of Wikimedia projects.

Nobody has ever thought search in Wikipedia or the various projects was particularly effective. However, bringing everything together doesn't just involve searching, and frankly there are a number of more pressing governance and community issues that need to be managed.

Perhaps I'm being a bit unfair here, but she was profiled when she first joined the WMF Board, and the following was said about her:

At the meeting, she described the impact on friends and family of the Chernobyl nuclear disaster, and the difficulty of getting reliable information in the face of “so much secrecy.”

Yet we see that this is precisely what happened with this grant proposal. A major grant was applied for and awarded and not even WMF staffers knew about it. You can see on the mailing list that it was a total shock when it was finally revealed.

I'm watching this train wreck from afar, but closer than others because some of my friends are deeply involved in Wikipedia and the WMF. I'm always amazed that a leadership change can complete kill an organisation. I've seen it in the corporate world, and I see it all the time in the volunteer world as well. The Wikimedia Foundation seems to be yet another victim of the appointment of a clueless leader, with no experience in the area or with the group they are meant to be leading, thrashing around, making changes without really understanding how systems work, the history of the organisation or relying on the experience and sage advise of the many expert and dedicated people around them, ultimately leading to a great deal of unnecessary turmoil, ill-will and frankly destruction in their wake.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#92

Earlier quoted context omitted.

That's OK, ironically it was Wikipedia that taught me to always back up my statements with references :-)

As an aside, I think the internet is simultaneously great at spreading absolute bs and disinformation and pushing people to have citations handy... paradoxically, both seem to be getting more frequent. (It all depends on where you browse, clearly).

Yeah, nobody knows this more than myself. I created [citation needed] and I've watched it be misused for years. I am glad I came up with the idea, but I'm resigned to the fact that it's human nature to misuse a valuable idea.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#93
post #38

Summary of the approach (p10): "1) Public curation mechanisms for quality; 2) Transparency, telling users exactly how the information originated; 3) Open data access to metadata, giving users the exact date source of the information; 4) Protected user privacy, with their searching protected by strict privacy controls; 5) No advertising, which assures the free flow of information and a complete separation from commerc…

[deleted]

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#94

Earlier quoted context omitted.

1) Public curation mechanisms for quality; The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc. 2) Transparency, telling users exactly how the inf…

Duckduckgo is ad free? I never knew this. How do they make money?

Ddg has an option to disable ads, they just ask that you help promote them.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#95
post #38

Summary of the approach (p10): "1) Public curation mechanisms for quality; 2) Transparency, telling users exactly how the information originated; 3) Open data access to metadata, giving users the exact date source of the information; 4) Protected user privacy, with their searching protected by strict privacy controls; 5) No advertising, which assures the free flow of information and a complete separation from commerc…

1) Public curation mechanisms for quality; The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc. 2) Transparency, telling users exactly how the inf…

> Curation doesn't scale and often assumes a single unifying ontology

Wikipedia is a pretty big exception to that assertion. Perhaps DMOZ (a clone of Yahoo circa 1996) is not the only way to do curation. Perhaps Wikipedia could apply what has worked for Wikipedia, i.e. develop a set of POV-neutral criteria for organizing collections of links and then invite everyone to participate.

It's really easy to be negative. But that's something that might at least be an interesting research project for the #1 open-curation system in the world.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#96
Good luck. We definitely need more search engines. (That Google announced to lower the PageRank(R)/site score for non-HTTPS sites is a clear indicator that they about to cross the line (monopoly). And no, DDG and most others are "just" meta search engines that rely on Yahoo Boss ($$$) which future is uncertain and relies on Bing.)

There was "Wikia Search" by Wikipedia founder Jimmy Wales:

"Wikia Search was a short-lived free and open-source Web search engine launched by Wikia, a for-profit wiki-hosting company founded in late 2004 by Jimmy Wales and Angela Beesley.

Wikia Search followed other experiments by Wikia into search engine technology and officially launched as a "public alpha" on January 7, 2008. The roll-out version of the search interface was widely criticized by reviewers in mainstream media. After failing to attract an audience, the site closed by 2009."

https://en.wikipedia.org/wiki/Wikia_Search

I used Wikia Search back then, it was good enough (like Bing in comparison to Google back then).

It was based on Apache Nutch and Solr/(Hadoop(?)/Lucene ...

Maybe you can rely on Lucene or SphinxSearch projects to kick-start.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#97

This is a very, very, very small amount of money if you want to build a search engine, let alone one "to rival Google" (source?). Looks like the goals are realistic, though - look how wikipedia search could be extended beyond results from wikipedia.org, build some test sets. And get a better idea what it really is that is supposed to be built.

Cuil spent about $30M before they went bust. By the time they went under, due to a total lack of a revenue model, they had a halfway decent search engine that did its own web crawl. So that's a data point on how much it costs.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#98

Earlier quoted context omitted.

I'd ask you to cite your claims, but we both know you can't. It's a pity your issues with Google cause you to pollute discussions with BS.

If you've ever read about the original PageRank algorithm, the parent post is a pretty reasonable way to describe it. I have no idea what the current algorithm looks like but I'd be shocked if it somehow switched to evaluating the 'quality' of content, however one might do that with an algorithm.

Well they do have some quality metrics, like duplication with other content, words used and so on. I suspect more, eg writing style measurements, correlated with other things that are found useful. There is a lot that could be done without actually understanding the content, although if course it can be gamed.

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#99
post #96

Good luck. We definitely need more search engines. (That Google announced to lower the PageRank(R)/site score for non-HTTPS sites is a clear indicator that they about to cross the line (monopoly). And no, DDG and most others are "just" meta search engines that rely on Yahoo Boss ($$$) which future is uncertain and relies on Bing.) There was "Wikia Search" by Wikipedia founder Jimmy Wales: " Wikia Search was a short-l…

Google doesn't want to lower the PageRank of http sites, it just wants to use http vs https as a feature in ranking. That isn't particularly surprising, I would be willing to bet google already uses hundreds of such features (one of which might be PageRank).

Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]

#100

This is a very, very, very small amount of money if you want to build a search engine, let alone one "to rival Google" (source?). Looks like the goals are realistic, though - look how wikipedia search could be extended beyond results from wikipedia.org, build some test sets. And get a better idea what it really is that is supposed to be built.

The basics of search are pretty well established today. Just because it initially cost Google a ton of money doesn't mean it would cost nearly as much today.

This holds true for nearly every human endeavor.

Post reply on HN