xMarks has an amazing set of data. A set of tagged bookmarks on a large enough scale is essentially like running mechanical turk (with 2,000,000 users) on the entire internet (better yet, only bookmark-worthy parts of the internet), and asking it to "describe this page in 10 words or less". The results, I think, are the most accurate description of URLs you could get on such a large scale. Essentially every site that…
Unfortunately, that dataset would only stay amazing until they tried to use it to power a search engine. History has shown that any time user-generated content is shown to users, spammers will quickly set themselves up as the ones generating 99% of the content.
Search is beside the point, though. They said they tried it and it was amazing but only for some things. It's true, just search delicious to get an idea of the type of search possible. It's very good for most things, but worthless for finding any content that people wouldn't tag, like 98% of the content of any URL.
My hypothesis was that their data could be used for other reasons.