Live data from Hacker News

Google's PageRank patent has expired (2019)

patents.google.com

101–110 of 173 posts

Re: Google's PageRank patent has expired (2019)

#101
post #75

Earlier quoted context omitted.

Search is objectively worse today than it used to be, it's not just that "it's harder to use for old people", it's just worse.

>Search is objectively worse today than it used to be How would you demonstrate that search is objectively worse? And how would you then show that it's a result of Google's algorithms, and not a consequence of the content of the Internet changing significantly?

> How would you demonstrate that search is objectively worse?

How would you demonstrate that being 50 years old is worse than being 25 years old?

You ask people that are 50 years old or older because they have been on both sides.

> it's a result of Google's algorithms

Well, it's simply in front of your eyes: this [1] was not possible in 2006.

Anyway, the fact that you cannot easily find on Google why Google search results are worse, proves that Google search results are worse today than in the past.

https://www.theverge.com/tldr/2020/1/23/21078343/google-ad-d...

Re: Google's PageRank patent has expired (2019)

#102
post #53

Earlier quoted context omitted.

Minor rant: I can't seem to contact an actual towing company directly, it's usually some external entity that ranks top in Google then they charge you more to find the actual local towing company. It says "Towing company in city name" but it's not local.

This was my experience last year when I wanted to find a locksmith. EVERY result on GMaps that was shown as being in my city was, in fact, some centralized company that seemed to contract out to guys working out of their cars. Every one took my name and said they'd get back to me (which they did). This is certainly a problem with the locksmith companies, but I think there's also a Maps problem, too: Google enables th…

Oh locksmiths are a fun one.

The type of people get into locksmithing are like those who do computer security for fun - they like to figure out how things work and how they can be broken. Which means that every locksmith wants to figure out how to game the system. And the first thing that they all figured out is that people tend to select whatever locksmith is closest. So they all went and pretended to be in a million places around the neighborhood in hopes that THEY would get selected.

That was the case a decade ago. And it was a nightmare. Glancing at a locksmith search now, Google somehow cleaned that up a lot. But it doesn't surprise me that whoever is on top now is someone who figured out how to game the current system.

Re: Google's PageRank patent has expired (2019)

#103
post #15

How did they even get that patent? It's a well known 70s era bibliometric algorithm.

The patent was actually assigned to Stanford University, and Google then licensed it (this is also how most pharmaceutical discovery is patented and licensed). A fundamental problem is that these patents all rely on federal funds from US taxpayers to one extent or the other, and while it may make sense for Stanford to hold the patent, there's a strong argument that any US entity should be able to license it (not just one exclusive license in other words).

Prior to Bayh-Dole legislation in the 1980s, this was the case for university-held patents: they could be not be exclusively licensed. Repealing that legislation would be a good idea to avoid the rise of monopolisitic behemoths like Google/Alphabet.

Re: Google's PageRank patent has expired (2019)

#104

Earlier quoted context omitted.

To be fair, they don't seem to teach the next generation about search operators. I've found just using quotes, OR/AND, and some other basic stuff can often get me what I need. The place where Google does well is that they're a monopoly, so things like image search that require a lot of resources, they are better at, especially paired with their geographic location. (USA! USA! USA!) Edit: I hit enter and the post went…

I've found search operators and quotes are less and less effective — especially on Google, but on other search engines as well — as time goes on. I think OR/AND were removed a long time ago, and things like + and quotes aren't effective, because they're still subjected to the same processing (stemming, synonym substitution, whatever ML nonsense Google does) as plain queries. So in general, you don't get what you're s…

>I've found search operators and quotes are less and less effective — especially on Google

On Google, I treat my search terms like a Venn diagram if I want good results.

Eg: want an article about the texas blackouts, but not the ones a decade ago?

Type "texas blackout npr 2022" minus quotes.

But if you do that on other engines, it may be MUCH more literal, the literal intersection of those terms, and I need to do the opposite: use as few terms as possible, possibly paired with using the site: operator, intitle operator, or other things.

>It's made it harder to search for bits of poetry, quotations, or song lyrics, especially.

Yeah to be completely clear, my default is DuckDuckGo, then very rarely I fall back to Google, but often if I'm doing that it's because I didn't want to trouble a librarian -- they talk about privacy, but I had a series of unfortunate events when I told one I want to use books as much as possible because I absolutely don't want some of these tech bros to know what I'm looking up.

(That dichotomy of folks who know information science and those who have critical thinking or coding skills needs to end, now. I'm an alumni of one of the highest ranked schools of information science in the world, and I will not be figuratively or literally extorted into a PhD to get roles others get with a bachelors.)

Re: Google's PageRank patent has expired (2019)

#105
post #97

Earlier quoted context omitted.

To be fair, they don't seem to teach the next generation about search operators. I've found just using quotes, OR/AND, and some other basic stuff can often get me what I need. The place where Google does well is that they're a monopoly, so things like image search that require a lot of resources, they are better at, especially paired with their geographic location. (USA! USA! USA!) Edit: I hit enter and the post went…

Try Yandex image search, you'll be surprised how good it is.

Yandex has an excellent reverse image search.

Re: Google's PageRank patent has expired (2019)

#107
post #36

Earlier quoted context omitted.

My search results were a lot better in 2006 when, I assume, they didn't have all these ML pipelines...

My view of it is that they basically lost the SEO spam wars. Without the ML pipelines, the top results would all be dominated by the smaller, highly skilled, SEO manipulators. They didn't find a way to cleanly excise that spam, so they resorted to a very imperfect hammer...giving a lot of SEO weight to large corporate entities. So basically, a different kind of spam dominates now.

There's a tonne of low hanging fruit google completely ignores, for example any page with an amazon referral link is almost certainly spam.

"But wait!", you say, "there are some legit reviewers out there." Yes, there sure are, by my starement is accurate, because for every legit review site with aws referrals, there are tens of thousands of ml created spam referral sites.

And so the real review sites are often lost in the mix regardless, which makss arguments to keep those results pointless.

But google leaves them there, and this is the same sort of site which, if it were an email, would immediately end up in a spam folder.

And beyond this, the other part of the problem is their ridiculous aliasing of search terms, which helps spammy sites come back as a response.

You say google lost? It's not losing, if you just don't care.

Frankly, it's just a return on investment thing. As long as only a few people per tens of thousands bolt, why spend the r&d?

Re: Google's PageRank patent has expired (2019)

#108
(Founder of Neeva here) Query-independent page and site signals like PageRank (or its site variants) have limited utility in a search ranker. Mostly, they are useful in weeding out bad pages and sites from your index, tie-breaking when you hit shard limits in retrieval and a few other edge cases.

The signals that matter the most: 1. Anchor text (and all variants of smearing and distinguishing between high quality and low quality anchors) 2. In aggregate, which pages got clicked on on any given query (and all smearing variants -- using ngrams, embeddings, ...). At Neeva, we use it for retrieval and scoring. 3. Query understanding signals mined from the query-click bi-partite graph and the query-query session refinement graph. 4. Page summarization signals built on top of 1 and 2 and body text. 5. To a lesser extent, query-independent page quality signals

Whether you use term-based retrieval or nearest neighbor (embedding) retrieval, a heuristic combination of signals or LambdaMart, whether you calibrate to human eval or clicks, whether your topical relevance function is hand crafted or uses a combination of deep learning and IR signals are all details past that.

tldr; there's a lot of craft in a search ranker, and no one silver bullet. Definitely not just PageRank.

Re: Google's PageRank patent has expired (2019)

#109
post #15

How did they even get that patent? It's a well known 70s era bibliometric algorithm.

The patent was actually assigned to Stanford University, and Google then licensed it (this is also how most pharmaceutical discovery is patented and licensed). A fundamental problem is that these patents all rely on federal funds from US taxpayers to one extent or the other, and while it may make sense for Stanford to hold the patent, there's a strong argument that any US entity should be able to license it (not just…

Doesn’t answer the question. Rephrased: Why was the patent granted if prior art existed from the 70s?

Re: Google's PageRank patent has expired (2019)

#110

Earlier quoted context omitted.

My search results were a lot better in 2006 when, I assume, they didn't have all these ML pipelines...

There was way less SEO back then. If they continued to use the old algorithm, your results would be worse than what you have now.

What are you saying?!

There was loads of SEO back in 2000 even! It brought alta vista to its knees, the number one search engine of the day.

Google got started, grew, because it filtered all that SEO spammy junk.

Do you have specific stats to back this up? Number of SEO pages vs good ones?

Or are you just presuming?

Post reply on HN