Live data from Hacker News

Google's PageRank patent has expired (2019)

patents.google.com

161–170 of 173 posts

Re: Google's PageRank patent has expired (2019)

#161
post #53

Earlier quoted context omitted.

Minor rant: I can't seem to contact an actual towing company directly, it's usually some external entity that ranks top in Google then they charge you more to find the actual local towing company. It says "Towing company in city name" but it's not local.

This was my experience last year when I wanted to find a locksmith. EVERY result on GMaps that was shown as being in my city was, in fact, some centralized company that seemed to contract out to guys working out of their cars. Every one took my name and said they'd get back to me (which they did). This is certainly a problem with the locksmith companies, but I think there's also a Maps problem, too: Google enables th…

If it's a Google Maps hit, street view and look for an actual storefront? Many mom-and-pop locksmiths have stores with safes etc on display. Though that won't work so well with tow companies.

Re: Google's PageRank patent has expired (2019)

#162
post #145
post #94

Earlier quoted context omitted.

> OK, but how is it harder for PageRank? If you are familiar with the algorithms, which I assume you are, you can work it out. To make my page score high on the PageRank score I need to acquire links from high PageRank score pages. This is a lot harder because it depends on a) in-links and b) high PageRank pages. With Hits, its easy for one page to harvest a high Hub score. All that is needed is to outlink to known g…

> With Hits, its easy for one page to harvest a high Hub score. All that is needed is to outlink to known good pages (authority). Providing outlinks is trivial. Once so harvested, one can direct that flow to a designated page to give it a high Authority score. Are you saying that you think HITS doesn't recursively score the quality of references by their own scores? That's not true. It does exactly what PageRank does…

> Are you saying that you think HITS doesn't recursively score ...

I doubt that reading comprehension is that hard a skill to master. I dont see where I have said anything about recursion or their lack of. If you want to have an imaginary conversation between yourself and what you think I have said, you can continue. I do not need to participate in that. I am sure you alone will suffice.

I think going back to the Pagerank and HITS papers carefully and understanding them will be illuminating. You keep saying they have no semantic difference, which cant be further from the truth. The scores are the eigenvectors of very different matrices, and HITS scores are straight forward to manipulate. BTW the papers cite each other stating in what way the other is different, if their difference was mere implementation detail and no semantic difference, I doubt they would stand as published papers.

The Hub score is not something interesting that I happened to say but a core part of the HITS paper. It is by virtue of the Hubs score that the Authority scores are defined (and vice versa) in the paper. It so happens, its easy to bump up the Hub score of a page by adding few strategic outlinks.

> Yes, again: possible to frame it as a formally valid problem if you really want to; still not an interesting one. We're only talking about this because you want to maintain that your earlier statement was true.

That's your opinion. I am just countering your categorical claim that there is no game theory formulation possible here. If your yardstick for your assertion that no game theoretic formulation is possible is your inability to find one that's interesting to you, there is not much I can do about it. All I can say is that such an yardstick is not very popular or useful.

A game theoretic formulation with a budget constraint adversary is a very natural setting. Research on link spam resistant node ranking algorithms are a thing, as is evaluating how stable are the rankings produced by some of these proposed algorithms to (potentially motivated and adversarial) changes to the links in the graph. SIGIR, WEBKDD proceedings on link analysis and rankings would be a good place to look.

You will need some background in matrix perturbation analysis, especially perturbation analysis of principal eigenvector to understand some of the results. Perturbation analysis of finite Markov chains will also suffice.

https://ai.stanford.edu/~ang/papers/ijcai01-linkanalysis.pdf

is by far one of the easiest papers to read in this area (Andrew Ng has focused on other areas of research since this paper). Note that the stability bounds in that paper can be easily tightened... left as an exercise for you. As you will see in the paper, HITS scores are easier to alter (equivalently stated, they are unstable) compared to Pagerank scores.

> Or maybe I'm wrong and there's a fascinating problem which you just don't want to divulge to me.

This is hardly the forum for extended discussions on a research topic. With the pointers and sketches that I mentioned a competent grad student would be able to fill in/ develop it further.

Re: Google's PageRank patent has expired (2019)

#163
post #25
post #14

Earlier quoted context omitted.

But does it make them less money?

"The goals of the advertising business model do not always correspond to providing quality search to users." http://infolab.stanford.edu/~backrub/google.html

And that's one of the major issues with google search: one metric which would be really good for a search engine is the inverse of the % of advertising related html / javascript in a page. Because that would minimize the utility of SEO for those link farms that only write text to mine ad views/clicks (like food recipes).

Alas, google would shoot themselves in the feet by promoting pages that consume less in their own advertising products.

If the US government split those two business units into different companies, we could have decent searches (google search could still profit by placing ads on its page results)

Re: Google's PageRank patent has expired (2019)

#164
post #150

Earlier quoted context omitted.

wikipedia? Beyond their remit for some queries for sure, but they fit the mold.

A search engine that only searched sites that wikipedia links to might be a fairly decent source in fact. If monetisable, it would turn gaming wikipedia into a whole new level of shitshow of course.

Mhmm great idea

Re: Google's PageRank patent has expired (2019)

#165

Earlier quoted context omitted.

Part of the problem is, I think, many searches strongly benefit from up to date content- programming tools, fashion, celebrities, things to do in X area, etc. It seems that Google has decided that most people want the most updated information when they look for something, which I don't think is entirely unreasonable. What I would love , however, is a way to turn that off for particular searches. Researching past even…

> What I would love, however, is a way to turn that off for particular searches But you can! In search results, click Tools and switch the Any time dropdown to Custom range... and you can specify a date range in the past. (Apparently, the custom option is hidden in the mobile version?!) I'm not sure how precise and dependable it is but it seems at least partially useful when I search for historical events.

I don't actually want to exclude new content, I just don't want to give it priority over older content if the older content is at least equally specific in matching my query.

Re: Google's PageRank patent has expired (2019)

#166

Earlier quoted context omitted.

You mention options in passing, but it is to me the root of the problem: Google hates giving up control. Control means ad revenue. So we could have options that would make search extremely efficient for most users, but that would presumably be very hard to monetize in comparison. So we have no options, and everyone gets mediocre to bad results. Since everyone I know in tech laments Google's decline into uselessness,…

> everyone gets mediocre to bad results Everything I've heard from people I know at Google suggests otherwise. Most searches for most people ... work. I too struggle to have google work in specific research cases, and I would like more power-user toggles, but basic searches like "$celeberty_name photos" or "$my_kids_school calendar" or "pizza places near me" just sorta work.

Hard to know anything for sure since results are "tailored". But Google used to be excellent for technical searches, whereas now it is unhelpful at the best of times. I'm guessing this isn't counted in "most searches gor most people".

Re: Google's PageRank patent has expired (2019)

#167
post #124
post #15

How did they even get that patent? It's a well known 70s era bibliometric algorithm.

> It's a well known 70s era bibliometric algorithm. Citation needed

I was thinking of the Pinski-Narin method, but there are other earlier uses of eigenvector calculations over very similar data: https://arxiv.org/pdf/1002.2858.pdf

Re: Google's PageRank patent has expired (2019)

#168
post #107

Earlier quoted context omitted.

There's a tonne of low hanging fruit google completely ignores, for example any page with an amazon referral link is almost certainly spam. "But wait!", you say, "there are some legit reviewers out there." Yes, there sure are, by my starement is accurate, because for every legit review site with aws referrals, there are tens of thousands of ml created spam referral sites. And so the real review sites are often lost i…

If there is a numbered list with referral links. 99% it's low effort referral spam trash. How they don't already filter this must be some sort active sabatague

Such links are relatively easy to hide from Google (obfuscate behind a redirect or inject with JS on user action), so if Google started using the links as a signal, spammers would hide the links, and invert the signal — only non-spammers who don't want to risk being delisted for cloaking would be left with the links, and penalized for them.

Re: Google's PageRank patent has expired (2019)

#169

Earlier quoted context omitted.

Whatever you do, whichever ML model you develop, a query-independent ordering of all documents will always be necessary since all distributed IR systems will over-retrieve according to some form of term matching and then apply sophisticated scoring to extract the best documents. You can't score trillions of docs for every query. Google still uses PageRank, but at the risk of stating the obvious, the current PageRank…

> You can't score trillions of docs for every query. That completely depends on how you model the queries. It can be a TB sized relation all the way into needing more bits than there are atoms on the universe.

This is not a generic discussion, we are discussing web search IR in particular, but if you want to be pedantic, the TB sized dataset could be the PageRank values and "scoring" could be the ordering of these values.

Re: Google's PageRank patent has expired (2019)

#170

Earlier quoted context omitted.

> You can't score trillions of docs for every query. That completely depends on how you model the queries. It can be a TB sized relation all the way into needing more bits than there are atoms on the universe.

This is not a generic discussion, we are discussing web search IR in particular, but if you want to be pedantic, the TB sized dataset could be the PageRank values and "scoring" could be the ordering of these values.

The TB dataset is PageRank indexed by search term. You can correlate them more and more and end up with exponentially more data, but even the smallest one (that you'll need a hack of servers to query) is quite useful already.
Post reply on HN