Earlier quoted context omitted.
1) Public curation mechanisms for quality; The Mozilla / open directory project tried this. Curation doesn't scale and often assumes a single unifying ontology. This is particularly problematic in a cross-cultural context. Besides, 'quality' is not a unidimensional metric in a result set: consider timeliness, authority, notability, uniqueness, comprehensibility, etc. 2) Transparency, telling users exactly how the inf…
Duckduckgo is ad free? I never knew this. How do they make money?
Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
81–90 of 192 posts
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#82Jimmy Wales' already tried to make a "Google Killer" ten years ago. It was tilting at windmills to say the least. Letting individuals help manage algorithmic search results was harder than you could imagine. Let's not even get into the difficulty of building an effective crawler.
One of Wikia's former CEOs, Gil Penchina, notoriously undervalued search as a result of this very public gaffe. By the time I came in, it took over five seconds to do a simple on-wiki search. Searching across wikis took so long they actually just sent the search to Google and had you abandon the site. I personally fixed a lot of these problems, and that part was pretty cool.
So now let's get to the subject at hand, which is a search feature based on an authoritative knowledge graph. Something like this should adequately surface factual information in an intuitive manner -- optimally based on natural language. Wikia already tried this, too. They brought on a very seasoned advisor who played a crucial role in the semantic web movement far back into the early oughts. I remember going to semantic web meetups in Austin when I was in grad school quite some time ago now to hear this guy talk.
This guy was essentially the SF-based manager or lead for a small team located in Poland whose job it was to take some of the "structured data" at Wikia and attempt to build some kind of knowledge graph on top of it. This project was unsuccessful.
So why did it fail? We'll start with a lack of product direction. Wikia had and probably still has a very junior product organization that is mostly interested in the site's UI and (recently) a focus on "fandom" (yuck). The team allocated to the project was based in Poland (Poznan, to be exact), and primarily kids coming out of a technical school on their first job. Your assumption about communication being a problem would be correct. However, the subject matter expert was so entrenched in his area of specialization, the problem was even more compounded on the native English-speaker side. There was too much getting in the weeds, and not enough focus on incremental progress.
To make things worse, they tried using a proprietary, not-ready-for-primetime data store because it most closely matched the SME's preconceptions on how the data should be structured. There was absolutely not an existing business use case for this data store, and problems getting it to work turned even building a simple demo into a death march.
Either way, what I'm saying is, $250,000 is not enough to solve this problem. We have attempted to solve this problem before in the MediaWiki world. It's not going to magically get better. To make something like this work, you need:
1) Best-in-class UX people who would know how a knowledge graph provides a significant improvement over existing solutions 2) Leadership that can bridge the gap between SMEs and implementers 3) Very skilled engineering resources with backgrounds in less conventional technologies
This is a massive investment that no one is willing to spend on what is essentially a media play.
About six months later, I had built a proof-of-concept that sucked data out of MediaWiki Infobox templates into Neo4j, a well supported graph database. I was able to answer questions like, "Which cartoon characters are rabbits", and "What movie won the most Oscars in 1968" using the Cypher query language.
At that point in time, Wikia had decided they were tired of investing in structured data, and wanted to re-skin the site for a third time in as many years to make it look more like BuzzFeed.
Structured data is cool. In many cases, unsupervised learning may be what you're actually looking for. But in the end it has to satisfy a real user's needs.
Wikipedia has five million English articles. Wikia has over 20 million. As far as capitalizing on this wealth of knowledge, the devil is truly in the details. But it's a real shame that all of that information isn't put to better use than to encourage the socially maladjusted to take quizzes over which anime character they're more like.
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#83Earlier quoted context omitted.
Can you say more/explain? What about the application is so upsetting? What is shameful about this? The Knight Foundation is about as upstanding as you can get, so it can't be that (full disclosure, I've received funding from them, so I'm definitely not unbiased on that point). So, what exactly is it that's so shameful here?
To be clear: my issue (and in fact, most peoples' issues) are not with the Knight Foundation. In fact, they appear to have been above board in every way in this whole debacle. It is the WMF board who are the problem here. See my comment here for just a few comments on this issue: https://news.ycombinator.com/item?id=11101262 Frankly, there's a lot more - to understand the issue better you might want to read Liam Wyat…
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#84Earlier quoted context omitted.
To be clear: my issue (and in fact, most peoples' issues) are not with the Knight Foundation. In fact, they appear to have been above board in every way in this whole debacle. It is the WMF board who are the problem here. See my comment here for just a few comments on this issue: https://news.ycombinator.com/item?id=11101262 Frankly, there's a lot more - to understand the issue better you might want to read Liam Wyat…
Thanks – I appreciate the references
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#85Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#86I'm not kidding when I say that if they want to know where to spend the $2.5m I would start with cleaning up their core codebase. IMO Mediawiki open source code is a disaster. EDIT: Not because it's written in PHP. Because it's architected poorly.
It's funny you should mention that. That was a point that apparently a number of WMF staff expressed, and it was apparently ignored.
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#87Earlier quoted context omitted.
Thanks – I appreciate the references
That's OK, ironically it was Wikipedia that taught me to always back up my statements with references :-)
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#88Earlier quoted context omitted.
I'm pretty sure that the spammer argument is just an excuse used by Google to allow them to keep their business practices out of public scrutiny. Google search results are biased in favour of content produced by those who have money and power. Google ranks everything based on popularity - Not based on quality. Popularity and quality are two independent concepts and not necessarily related. That's something which Wiki…
I'd ask you to cite your claims, but we both know you can't. It's a pity your issues with Google cause you to pollute discussions with BS.
I have no idea what the current algorithm looks like but I'd be shocked if it somehow switched to evaluating the 'quality' of content, however one might do that with an algorithm.
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#89Summary of the approach (p10): "1) Public curation mechanisms for quality; 2) Transparency, telling users exactly how the information originated; 3) Open data access to metadata, giving users the exact date source of the information; 4) Protected user privacy, with their searching protected by strict privacy controls; 5) No advertising, which assures the free flow of information and a complete separation from commerc…
I'm pretty sure that the spammer argument is just an excuse used by Google to allow them to keep their business practices out of public scrutiny. Google search results are biased in favour of content produced by those who have money and power. Google ranks everything based on popularity - Not based on quality. Popularity and quality are two independent concepts and not necessarily related. That's something which Wiki…
Re: Wikipedia starts work on $2.5M internet search engine project to rival Google [pdf]
#90I am guessing this has a different focus than their previous attempt at making a search engine, wikia search, which they abandoned fairly quickly https://en.wikipedia.org/wiki/Wikia_Search
https://en.wikipedia.org/wiki/Wikia#Relationship_with_Wikipe...
http://community.wikia.com/wiki/Help:Wikimedia
(Edit: Who decided that enter is not equal to enter?)