Live data from Hacker News

Google Web spam - Gabriel Weinberg's Blog

gabrielweinberg.com

21–23 of 23 posts

Re: Google Web spam - Gabriel Weinberg's Blog

#21

Earlier quoted context omitted.

They mean something, i.e. that they "are in their index in some form." I've been blacklisted before, and when you're blacklisted, you don't show up in site: queries. That said, I wanted to acknowledge that this isn't ranking data. However, perhaps as a result of this post, I'll be able to get some and re-post those results.

I'm sure you would agree with this but in case others are reading, simply blacklisting these sites wouldn't be the best thing to do. Many are simply expired or parked pages.

Google visits domains all the time so they should be aware quite quickly when things move from spam/parked to non-spam/non-parked. Therefore, I don't see why they shouldn't all be out of the index until they have useful content on them.

Re: Google Web spam - Gabriel Weinberg's Blog

#22

Earlier quoted context omitted.

Can you please explain "impression-weighted precision" and how it's used? Sounds very interesting.

In general we want to measuring what matters to users, so removing a lot of spam that nobody sees doesn't really make anything better. By the same token, if you are trying to launch a new spam classifier that has some false positives, if one of the false positives is yahoo or facebook, it doesn't really matter how good it is, it will never be worth the collateral damage. As a result, rather than measuring precision/r…

"As a result, rather than measuring precision/recall as a percentage of domains or as a percentage of urls we usually try to measure it as a percentage of results that appeared on a search result page or results that users click on, mined from our logs."

So this basically means that you are able to discern content spam on authoritative domains (facebook, wordpress.com, etc) based on ctr, bounce, impressions compared to surrounding serp results rather than comparing that data against the parent domain as a whole?

Re: Google Web spam - Gabriel Weinberg's Blog

#23

Earlier quoted context omitted.

In general we want to measuring what matters to users, so removing a lot of spam that nobody sees doesn't really make anything better. By the same token, if you are trying to launch a new spam classifier that has some false positives, if one of the false positives is yahoo or facebook, it doesn't really matter how good it is, it will never be worth the collateral damage. As a result, rather than measuring precision/r…

"As a result, rather than measuring precision/recall as a percentage of domains or as a percentage of urls we usually try to measure it as a percentage of results that appeared on a search result page or results that users click on, mined from our logs." So this basically means that you are able to discern content spam on authoritative domains (facebook, wordpress.com, etc) based on ctr, bounce, impressions compared…

No, not at all. This is just talking about the granularity at which we compute metrics like "spam rate."
Post reply on HN