Live data from Hacker News

Google Web spam - Gabriel Weinberg's Blog

gabrielweinberg.com

11–20 of 23 posts

Re: Google Web spam - Gabriel Weinberg's Blog

#11
post #9

Being in the index isn't very meaningful. Very little of our spam-fighting/ranking prevents sites from showing up for "site:" queries, because in general we think that if someone is intending on going to a domain directly, the only reasonable thing to do is to show that domain. Unfortunately I don't have a better way of assessing it to offer. Internally, we often look at impression-weighted precision as a metric, but…

I think this method would violate (at least the spirit of) my privacy policy: http://duckduckgo.com/privacy.html That said, if anyone has a meaningful sample query set, I'm certainly interested in running it against my spam index. I see a lot of hits on it via other search APIs.

You could let people opt their queries into such studies.

Re: Google Web spam - Gabriel Weinberg's Blog

#12
post #11

Earlier quoted context omitted.

I think this method would violate (at least the spirit of) my privacy policy: http://duckduckgo.com/privacy.html That said, if anyone has a meaningful sample query set, I'm certainly interested in running it against my spam index. I see a lot of hits on it via other search APIs.

You could let people opt their queries into such studies.

I don't have accounts right now so it's a bit tricky. If there was some way people could export their Google search history, then I could use that. Feel free to email it in anonymously.

Re: Google Web spam - Gabriel Weinberg's Blog

#13
post #9

Being in the index isn't very meaningful. Very little of our spam-fighting/ranking prevents sites from showing up for "site:" queries, because in general we think that if someone is intending on going to a domain directly, the only reasonable thing to do is to show that domain. Unfortunately I don't have a better way of assessing it to offer. Internally, we often look at impression-weighted precision as a metric, but…

Can you please explain "impression-weighted precision" and how it's used? Sounds very interesting.

Re: Google Web spam - Gabriel Weinberg's Blog

#14
post #9

Being in the index isn't very meaningful. Very little of our spam-fighting/ranking prevents sites from showing up for "site:" queries, because in general we think that if someone is intending on going to a domain directly, the only reasonable thing to do is to show that domain. Unfortunately I don't have a better way of assessing it to offer. Internally, we often look at impression-weighted precision as a metric, but…

Can you please explain "impression-weighted precision" and how it's used? Sounds very interesting.

In general we want to measuring what matters to users, so removing a lot of spam that nobody sees doesn't really make anything better.

By the same token, if you are trying to launch a new spam classifier that has some false positives, if one of the false positives is yahoo or facebook, it doesn't really matter how good it is, it will never be worth the collateral damage.

As a result, rather than measuring precision/recall as a percentage of domains or as a percentage of urls we usually try to measure it as a percentage of results that appeared on a search result page or results that users click on, mined from our logs.

This is one of the bajillion reasons why it's absurd to expect Google to throw away all its logs data. The logs are essential to coming to any meaningful conclusions about the current quality of our search, let alone finding ways to improve it.

Re: Google Web spam - Gabriel Weinberg's Blog

#15

Earlier quoted context omitted.

Can you please explain "impression-weighted precision" and how it's used? Sounds very interesting.

In general we want to measuring what matters to users, so removing a lot of spam that nobody sees doesn't really make anything better. By the same token, if you are trying to launch a new spam classifier that has some false positives, if one of the false positives is yahoo or facebook, it doesn't really matter how good it is, it will never be worth the collateral damage. As a result, rather than measuring precision/r…

Thanks.

Am I right in understanding that if a SERP result for a given keyword doesn't get clicked by users enough, it will be removed?

By the same token: does the result's bounce rate matter? I imagine spammy sites have a very high bounce rate.

Re: Google Web spam - Gabriel Weinberg's Blog

#16
post #11

Earlier quoted context omitted.

You could let people opt their queries into such studies.

I don't have accounts right now so it's a bit tricky. If there was some way people could export their Google search history, then I could use that. Feel free to email it in anonymously.

I found this link which looks like it dumps search history as an rss feed. That might be a convenient way to send it over.

https://www.google.com/history/?lookup?q=&output=rss&#38...

Re: Google Web spam - Gabriel Weinberg's Blog

#17
post #5
post #2

On an tangential note, it's a shame that Google don't offer a service like BOSS (though BOSS could be enhanced by offering revenue sharing, as an alternative to charging per query). It seems the Google search APIs have actually gone backwards over the last few years.

Yahoo's goal with BOSS is to fragment the search market; Google is so far ahead that Yahoo knows they can't compete head-on. Google doesn't want the search market to be fragmented; they want to dominate the market. I think that's why Google doesn't have a good search API offering.

The actual reason for this is that Google is extremely (unreasonably?) paranoid about leaking how we do things to our competitors.

Re: Google Web spam - Gabriel Weinberg's Blog

#18

Earlier quoted context omitted.

In general we want to measuring what matters to users, so removing a lot of spam that nobody sees doesn't really make anything better. By the same token, if you are trying to launch a new spam classifier that has some false positives, if one of the false positives is yahoo or facebook, it doesn't really matter how good it is, it will never be worth the collateral damage. As a result, rather than measuring precision/r…

Thanks. Am I right in understanding that if a SERP result for a given keyword doesn't get clicked by users enough, it will be removed? By the same token: does the result's bounce rate matter? I imagine spammy sites have a very high bounce rate.

[deleted]

Re: Google Web spam - Gabriel Weinberg's Blog

#19

Earlier quoted context omitted.

In general we want to measuring what matters to users, so removing a lot of spam that nobody sees doesn't really make anything better. By the same token, if you are trying to launch a new spam classifier that has some false positives, if one of the false positives is yahoo or facebook, it doesn't really matter how good it is, it will never be worth the collateral damage. As a result, rather than measuring precision/r…

Thanks. Am I right in understanding that if a SERP result for a given keyword doesn't get clicked by users enough, it will be removed? By the same token: does the result's bounce rate matter? I imagine spammy sites have a very high bounce rate.

One reason why spam sites have high bounce rates is because they act as funnels to other websites.

Re: Google Web spam - Gabriel Weinberg's Blog

#20
post #3

The writer states himself that the results of "site:" don't mean anything: "Of course this says nothing about how much they appear in the rankings." So what's the point of this article?

They mean something, i.e. that they "are in their index in some form." I've been blacklisted before, and when you're blacklisted, you don't show up in site: queries. That said, I wanted to acknowledge that this isn't ranking data. However, perhaps as a result of this post, I'll be able to get some and re-post those results.

I'm sure you would agree with this but in case others are reading, simply blacklisting these sites wouldn't be the best thing to do. Many are simply expired or parked pages.
Post reply on HN