Live data from Hacker News

How My Popular Site was Banned by Google

kbrower.posterous.com

101–110 of 132 posts

Re: How My Popular Site was Banned by Google

#101
post #69

Earlier quoted context omitted.

Literally Google Webmaster Guidelines say: "Use robots.txt to prevent crawling of search results pages or other auto-generated pages that don't add much value for users coming from search engines." So a couple of questions: - What is of value for a user? - Who is determining the value for the users in these cases? As always, it's not always quite clear on what treatment you should use for search pages!

I'm pretty sure that having different search engines indexing each other would be bad. Don't cross the streams.

If I recall correctly, crossing the streams did kill marshmallow man!

In any case, I agree with you, search engine indexing search results would be bad, but the line is not that clear all the time!

Some vertical search engine result pages are a great and relevant result from a user perspective on the question they are trying to solve.

Re: How My Popular Site was Banned by Google

#105
post #14

From google's bridge page definition: Not acceptable: "Websites that feature links to other websites while providing minimal or no added functionality or unique content for the user Added functionality includes, but isn't limited to, searching, sorting, comparing, ranking, and filtering" The page in question provides the added functionality listed (searching/filtering at least)

The filleritem.com website is a very simple tool to search for items of a certain price on Amazon. The purpose is to help select an item to fill the gap in the shopping cart to qualify for free shipping. There is no reason for Google to ban this site. What we MAY have here is an abuse of power by Google.

Re: How My Popular Site was Banned by Google

#107

Earlier quoted context omitted.

As we have heard lately how Google has manipulated the Yelp search results and others for their benefit, one has to wonder - since google has aquired some big websites like Zagat, theoretically - if one day they will just start to screw Zagat competitors , what can be actually done about it , i mean how can one prove that google is just attacking you and they just dont like you as their competitor ... its relative -…

It's a fair question, but our quality guidelines have been quite consistent for the last decade or so. Our team takes action without regard to whether a website is an advertiser, partner, or competitor. The long-term loyalty/trust of our users is worth much more than any sort of short-term revenue that a sneaky deal would provide. And Google's culture is such that anyone inside or outside of the company can claim tha…

Matt: This is a lie. I have seen you talk about quality at SMX and it is laughable. These quality guidelines you speak of might pertain to others, but they do not pertain to Google's own content.

1. When I search for a restaurant, 99% of the time Google places shows up before Yelp. So you are telling me that Google places always has better content than Yelp? Look at these reviews for Gary Danko - http://maps.google.com/maps/place?num=100&hl=en&biw=...

"love it" and "our favorite restaurant" Does Google Places have much better quality content 99% of the time?

2. Look at http://www.seobook.com/images/google-google-google-google-go.... A. Look at all of Google's sites in the results. B. What are those Youtube videos doing in these results? Every other result in this set has the words hollywood and cauldron right together(except the leaky-cauldron domain result), but since Youtube is a Google property it shows up. It's funny. Do this search on your own and you will see that the top 40 results have "Hollywood Cauldron" in its title, but the ones in Youtube do not.

So your results are all about quality unless it comes down to your own content. Then, it doesn't matter.

By the way, I can't wait for Google Plus to start hogging space in your results too. We all know how important "social" is to you guys. Even more important than "local." So we shall assume the results will be even more crowded google content. If we are lucky, maybe Robert Scoble will mention the words hollywood and cauldron in one of his wordy posts and we will see at the top of the results! And by the way, no one buys that don't be evil thing any more... long gone.

Re: How My Popular Site was Banned by Google

#108

Earlier quoted context omitted.

This is a fair comment but I believe the issues stem more from the process Google follows. It is obvious that Google cannot communicate exact reasons why a site was penalized as that would help spammers. However, there is nothing that prevents them from adding a step to warn the offending website and give them a heads up before the ban/penalty takes place, along with an explanation of the policy that is/was being vio…

option1138, I'm afraid you need to recalibrate your expectations of spam on the web. Blekko made a site called Spam Clock that estimates 1 million spam pages are created every hour : http://www.spamclock.com/ . There's 200+ million websites out there. 1,000 spam sites would be a spam rate of 5.0 × 10^-6. If you remember the days of Altavista before Google, the actual rate of spam on the web is much higher. Here's one…

Matt,

You are of course correct. The fault is mine for miscommunicating... I find myself becoming less self-editorial these days when I write on the web and tend to think everyone is on the same page as I am.

I was actually referring to an informal study I did earlier this year. I measured sites which were receiving an average of 50,000 or more visitors from Google US search (organic) per month over a six month period. Then I compared those with a similar set from a subsequent six month period to see which had significantly dropped off in traffic and rankings. The purpose of this was to estimate the number of significant sites which were penalized over that period of time. The final estimate came to about 700 sites/year which were penalized. There are lots of uncontrolled variables here of course... but I was looking for an "order of magnitude" answer simply for curiosity's sake.

The 1 million spam pages created per day were of course excluded from consideration as they never received much traffic from Google in the first place.

So just to clarify my earlier response, I am advocating for a policy that would apply to websites exceeding a certain threshold of organic traffic for a significant period of time.

Re: How My Popular Site was Banned by Google

#109

Earlier quoted context omitted.

Sadly, quite true. "True" - because my current understanding (which Matt_Cutts can elucidate on if he chooses to) is that Google has looked into - but does not currently incorporate - the presence of advertising as a spam signal. "Sadly" - because my independent research has shown that advertising - most notably the presence of Google AdSense - is a reliable predictive variable of a page being spam. All things being…

I agree but don't you think that from an algorithmic point of view Google would be better looking at what the user wants and what monetization models they prefer to see versus the averages in terms of monetization models on spam sites. That way their focussing less on removing spammers and more on user quality and thus removing spam.

This really made me think. But I took exception to your comment:

I agree but don't you think that from an algorithmic point of view Google would be better looking at what the user wants and what monetization models they prefer[...]

No. In fact, I am a rather loudmouthed opponent to Google's somewhat clumsy attempts to measure this ala "Quality Score".

In addition to webspam detection and machine learning, I have spent way too much time in marketing (I have a master's degree in marketing, in fact.)

A neat thing I learned along the way was the value of market research.

There are so many nuances in every line of business. Segments, preferences, pricing, even down to minutia (now well studied) such as fonts, gutter widths, copy styles, and so on.

You can learn a lot by combining large amounts of data and well chosen machine learning algos. But even with a few thousand businesses in most categories in a particular country (far less outside of the US), that doesn't give an outsider enough data to truly distinguish what can be a winning formula from a spammy one. This knowledge is hard won through carefully executed experiments and research.

A few years ago I was researching the topic of landing page formulas by category. One example that stuck out most in my mind was mortgages. There were a few tried and true "formulas" that significantly outperformed the rest. Two stuck out:

1) Man, woman, and sometimes child standing on a green lawn in front of home. Arrow pointing down from top left of landing page to mid/lower right positioned form. Form limited to three fields.

Edit: http://imgur.com/90VmB

2) Picture of home/s docked to bottom of lead gen page. No people. Light/white background. Arrow pointing down from top left of landing page to centrally located form.

Edit: http://imgur.com/JkLlH

These sites were incredibly successful. More than a few of them had to contend with quality score issues over the years. Can an algorithm capture nuances such as the ones I mentioned? In theory... they could. But today, they don't. All of Google's QS algorithms to date have been failed attempts and have caused an incredible amount of harm and distrust.

You finished that sentence with:

to see versus the averages in terms of monetization models on spam sites.

I'm not at all sure what this means. Could you explain? Is it even possible to directly model the monetization model of a site without having direct access to their metrics?

Re: How My Popular Site was Banned by Google

#110

Earlier quoted context omitted.

I recently talked for about a minute about the topic of "too much advertising" that sometimes drowns out the content of a page. It was in this question and answer session that we streamed live on YouTube: http://www.youtube.com/watch?v=R7Yv6DzHBvE#t=19m25s . The short answer is that we have been looking more at this. We care when a page creates a low-quality or frustrating user experience, regardless of whether it's…

Thanks Matt, took a look at the video, that answers my original question. Also is there a preferred monetization model, e.g. do Google think advertisements are more or less harmful to the user experience than affiliate links, sponsored posts, etc? Obviously across different models you can't just track space taken up, so is there some kind of metric that tracks the rate of content diluting via monetization?

I know you were looking for an answer from Matt but I thought I would offer up my opinion here (as you might have noticed, I love talking about this stuff).

Our models currently suggest that the presence of contextual advertising is a significant predictive factor of webspam.

We use 10-fold bagging and classification trees, so it's not all that easy to generalize. But I pulled one model out at random for fun.

The top predictive factor in this particular model is the probability outcome of the bigrams (word pairs) extracted the visible text on the page. Here are a few significant bigrams:

relocationcompanies productsproviding productspure qualitybook recruitmentwebsite ticketsour thesetraffic representingclients todayplay tourshigh registryrepair rentproperties weddingportal printingcanvas prhuman privacyprotection providingefficient waytrade printingstationery priceseverything website*daily

Next, this model looks for tokens extracted from the URL and particular meta tags from the page. Similar to above, but I believe unigrams only. A few examples follow. Please keep in mind that none of these phrases are used individually... they are each weighted and combined with all other known factors on the page:

offer review book more Management into Web Library blog Joomla forums

The model then looks at the outdegree of the page (number of unique domains pointed to).

From there, it breaks down into TLD (.biz, .ru, .gov, etc)

The file gets pretty hard to decipher at this point (it's a huge XML file) but contextual advertising is used as a predictive variable throughout.

Just from eyeballing it, it appears to be more or less as significant as the precision and recall rate of high value commercial terms, average word length (western languages only), and visible text length.

Based on what I'm looking at right now, my answer would be that sponsored posts are going to be far more harmful to the user experience than advertising.

Can't answer the rest of your question which I assume relates to the number of ad blocks or amount of space taken up by ads... we don't measure it.

Edit: Just realized that Google will probably delist this page within 24 hours. Should've used a gif for those bigrams. Oh well ;-)

Post reply on HN