Live data from Hacker News

The Age of PageRank Is Over

blog.kagi.com

241–250 of 380 posts

Re: The Age of PageRank Is Over

#241
post #164

Earlier quoted context omitted.

For a long time, I've been considering a different solution to the problem, which is to create a human-curated whitelist-only search engine. I do think your idea is great and something with a lot of value for many use cases. The only issue is that it won't surface certain things like restaurant contact info and Wikipedia that people often need to be a top result.

>which is to create a human-curated whitelist-only search engine. You might be interested in watching the series "Halt and Catch Fire".

I'm not interested. Why would I be?

I also don't need fiction to know what a human-curated search engine would be like, because I was using the web in the 90s when we had things like Dmoz.

Re: The Age of PageRank Is Over

#242
post #175

Earlier quoted context omitted.

"a human-curated whitelist-only search engine" This is a good idea. Basically harkening back to the old internet directory days, combined with the powerful indexing tools we have now, and then use all the great language models and topic modeling tooling to make the querying great. Said another way: moderate and protect the index, instead of trying to clean up the results of queries. You'd need to monitor the whitelis…

> combined with the powerful indexing tools we have now What powerful indexing tools exist today? I have little context on the search domain but it's very interesting.

I'm not sure, but it sounds like they mean that building and quickly searching an index has been democratized. You can now use any number of FOSS projects or SaaS offerings to do that, depending on how big you want the index to be and how much you want to spend.

Re: The Age of PageRank Is Over

#243
Google is not using raw PageRank anymore, they use some sort of machine learning, which is overly optimized based on flawed abusable SEO techniques and probably user count, rather than actual value and authenticity.

Re: The Age of PageRank Is Over

#244

IMO this article misses the biggest issue with search engines today. It's less about any ranking algorithm (like PageRank), but rather the root cause is with indexing. No matter how much you fine tune your ranking system, if your index is filled with SEO junk and blogspam, you're fighting a losing battle from the start. For some reason, a lot of these search engines like to brag about the number of documents in their…

Kagi has a feature, lenses, which includes a discussion lens. It’s pretty much that, results from forums and reddit.

Re: The Age of PageRank Is Over

#245

Earlier quoted context omitted.

No, I think a modern directory would be a set of topics in an ontology with links. If I had to seed one I would suck all the external links out of Wikipedia, impose some organizing structures (probably overlapping trees or dags) and then build a set of classifiers for nested topic relevance, spaminess, etc. There are certain sites which have a landing page for every topic in some set of topics, you could add a lot of…

Exactly how do you think Google (and Bing, and ...) work? They do start from known indexes, in particular wikipedia. Hell, Google even has internal papers where they claim they improved search quality by "wikipedia's" as a unit.

Actually Microsoft bought a company called Powerset which did information from Wikipedia to build a "semantic as in semantic web" index, that technology became the heart of Bing.

Google was caught flat footed and wound up buying and killing Freebase in order to catch up. They lied and said they rejected "semantic web" approaches despite hiring one of the leaders of the Cyc project as their head of research.

Still there is a difference between exposing that kind of database through a full text index vs exposing it through a browsing interface.

Re: The Age of PageRank Is Over

#246
post #175
post #164

Earlier quoted context omitted.

For a long time, I've been considering a different solution to the problem, which is to create a human-curated whitelist-only search engine. I do think your idea is great and something with a lot of value for many use cases. The only issue is that it won't surface certain things like restaurant contact info and Wikipedia that people often need to be a top result.

"a human-curated whitelist-only search engine" This is a good idea. Basically harkening back to the old internet directory days, combined with the powerful indexing tools we have now, and then use all the great language models and topic modeling tooling to make the querying great. Said another way: moderate and protect the index, instead of trying to clean up the results of queries. You'd need to monitor the whitelis…

> You might be able to jump start the whitelist by getting folks using a plugin so you can know what domains they spend time on in the first place and index those when signal strength warrants it.

I actually think a browser plugin is a core component, especially if people could opt in to contributing their pages to the index (with a method to make sure they're not logged in to anything).

The plugin would also allow people to flag websites as spam/scams/low-value, regardless of whether they're in the index or not.

Re: The Age of PageRank Is Over

#247
post #165

I've been using Kagi the last few months. I've never had a reason to complain about its search results - they seem to work plenty well enough day to day that I don't really 'notice' them. Like most, I now search reflexively, as an extension of the mind, so I only 'notice' search when it's bad. What really excites me is that is that I'm paying them. That sounds odd, but seriously. It's incredibly refreshing to know th…

>It's incredibly refreshing to know that the company providing my search results has an incentive to make things better for _me_ and not a legion of advertisers Cable television and netflix have made it quite clear that payment does not mean “no ads”, it just means “no ads, yet”. It still might be a better alternative in the short term, but the moment growth starts to peak, companies follow where the incentives lead…

Can you add a poison pill to a company where a select group of people can decide you broke one of your founding principles and hand the company over to someone else?

No VC startup would do that but if possible would allow trust.

Re: The Age of PageRank Is Over

#248

In a word, no. I’ve spent a lot of time on SEO over the past couple years, and inbound links still matter a lot to search rankings and traffic. This is clear evidence that PageRank still matters. From a more macro perspective, I’ll believe Google is failing when a competitor starts eating their lunch. What I see right now are a bunch of would-be competitors who want to eat their lunch, including this company. The blo…

I think they're more like trying to grab several of the crumbs that Google left behind. I don't think it's possible for a paid search engine to eat Google's lunch.

Yeah, the devs stated again and again, that they don’t plan or expect to be a mainstream search engine.

Re: The Age of PageRank Is Over

#249
I don't understand why advertising-based revenue models are bad for consumers. The most profitable ads are those that are the most relevant to the consumer. It leads to advertisements that are most likely to provide commercial information that the consumer finds valuable.

Re: The Age of PageRank Is Over

#250
post #164

Earlier quoted context omitted.

For a long time, I've been considering a different solution to the problem, which is to create a human-curated whitelist-only search engine. I do think your idea is great and something with a lot of value for many use cases. The only issue is that it won't surface certain things like restaurant contact info and Wikipedia that people often need to be a top result.

I like this idea in theory, but I think it would be difficult to scale up enough to be useful. Who would be the curators? If you open up curation to volunteers, it could be gamed by bad actors. However if you have too small a team, the results will be limited and biased. For example, favoring the English-speaking or tech-sphere web while ignoring large sections of the web with which the curators are unfamiliar. Perha…

> Who would be the curators?

That's a key question. You don't necessarily need that many of them, though, if you're whitelisting at the domain level. A few dozen people could work giving a rating to many thousands of domains every year. An ideal number might be in the thousands.

I'd also say they probably shouldn't be in industry, because they'd have an incentive to game the index.

> Perhaps machine learning could help with scale - start with a human curated dataset, then train a model on it.

While this is true, I don't think search engines other than Google and Bing need to worry that much about being gamed. It's just not worth the effort.

Post reply on HN