Earlier quoted context omitted.
I talked to someone from Google at I/O who should know and he claimed they don't play "Whack a mole" with websites. They will tweak their ranking algorithm to punish the behavior they see in a web site they don't want to be ranked.
That was probably me. We have two sides to the webspam team at Google: engineering and manual. We definitely prefer to write algorithms so that we avoid dealing with individual websites--the idea is that you strive to fix the root cause of an issue, not to tackle specific sites. However, if we see a website that violates our guidelines and that gets past the algorithms, we are willing to take manual action. Where pos…
My history of (mostly failed) side projects and startups
61–70 of 82 posts
Re: My history of (mostly failed) side projects and startups
#62Why did Google blacklist all of your Tldscan sites? Was it just because your sites' content was updated automatically? Or was it because you did something wrong for SEO?
I have a few websites that automatically make new posts. As of 10/14, they all show 0 pages indexed in Google. Previously they would get a few thousand visitors per day.
I guess Google feels as though they violate their terms and removed them. It seems to me it was a manual removal.
I received no emails in webmaster tools about the removal.
Re: My history of (mostly failed) side projects and startups
#63Earlier quoted context omitted.
From what I've seen Google doesn't contact people :) My guess is they also have a policy of not sharing reasons for getting blacklisted, to ensure they're not giving spammers an easy way to fix their website. They claim they respond to all "Site reconsideration" requests. I had to file one once, they did respond, but with a very non-informative and unhelpful response.
Yeah, in retrospect I should have taken it slower and not gotten as close to the line in the first place. It's totally my fault, and I'm not bitter. As you can tell from the OP, I've had a lot of failure, and I similarly learned from this one.
The net effect is that we haven't found a way to talk 1:1 with every webmaster, and I'm not sure whether that's possible. The story of webmaster communication for the last few years at Google has been trying to improve scalability of the info. The earliest Google webmaster communicator ("GoogleGuy") answered questions on a webmaster forum. In 2005 I started a blog, which has the advantage of permalinks for posts like http://www.mattcutts.com/blog/seo-mistakes-autogenerated-doo... . We tried doing live webmaster chats, but that would only reach 400-500 webmasters at a time.
The most scalable thing I've found so far is making videos. Here's a video that came out last month about the dangers of autogenerating pages for example: http://www.youtube.com/watch?v=A8bgpWtVHo4 . We're at almost 300 videos now, and we're getting closer to 3M total views on our webmaster video channel. The hope is that this additional guidance helps people self-identify what can cause issues to avoid or to correct them without needing to talk to Google.
The other big tool that has been helpful is http://google.com/webmasters/ . That provides tools to identify the common errors/mistakes that webmasters make (crawl errors, 404 pages, canonicalization, robots.txt issues, identifying hacked sites using the "Fetch as Googlebot" feature, etc.). That helps with many of the straightforward issues, but of course it doesn't solve the issue with "sheer number of webmasters who have ranking questions vs. number of Googlers." If anyone has suggestions on how to tackle communication with webmasters in a more scalable way, I'd appreciate feedback on how to do better on that.
Re: My history of (mostly failed) side projects and startups
#64Earlier quoted context omitted.
Yeah, in retrospect I should have taken it slower and not gotten as close to the line in the first place. It's totally my fault, and I'm not bitter. As you can tell from the OP, I've had a lot of failure, and I similarly learned from this one.
The tricky part is that the math works out something along the lines of there being ~200,000,000 domains and there being ~20,000 Google employees. At a simplistic level that works out to 10,000 domains per Google employee. Which means that even if Google stopped doing everything else and everyone at Google spent all their time talking to webmasters, they'd each have to answer 10,000 peoples' questions about rankings,…
Second, we're doing our part to spread what we're learning about running a top 500 website (Stack Overflow) with the community, in the form of http://webmasters.stackexchange.com
Do we make mistakes? You bet we do. Just the other day I accidentally disallowed all questions on Stack Overflow from being spidered in robots.txt. That.. was .. not a good day.
Re: My history of (mostly failed) side projects and startups
#65As a builder of digital things, the 3 things I have the most trouble communicating to non-builders are: - how hard it is - how much time it takes - how long it takes to become successful So instead of trying to explain it, I may just send them to this blog post, which shows all 3. Thank you, Gabriel! (Now if only you would remove that Mojo Badge business from blocking your great content.)
Re: My history of (mostly failed) side projects and startups
#66Earlier quoted context omitted.
Yeah, in retrospect I should have taken it slower and not gotten as close to the line in the first place. It's totally my fault, and I'm not bitter. As you can tell from the OP, I've had a lot of failure, and I similarly learned from this one.
The tricky part is that the math works out something along the lines of there being ~200,000,000 domains and there being ~20,000 Google employees. At a simplistic level that works out to 10,000 domains per Google employee. Which means that even if Google stopped doing everything else and everyone at Google spent all their time talking to webmasters, they'd each have to answer 10,000 peoples' questions about rankings,…
I understand the argument behind keeping it a black box, but it doesn't need to be as much of a blackhole. For example, in this case the following could have happened:
1) Site triggers some alarm for violating something.
2) Just those site(s) get strongly penalized.
3) Automatic emails go out in the message centers of Google Webmaster tools, analytics, adsense, and Gmail -- wherever the sites show up registered. In my case, it would have been all of the above.
4) The messages indicate the nature of the violation, that there is a penalty in effect.
5) There is a link to click on if you think you've corrected the errors.
6) If you click it, it auto-checks your site in y days and sends you another message that it passed or not.
7) If not corrected, it stays penalized or there are a series of penalties until full blacklisting.
That's all automated, i.e. scalable. I understand there are some tricky bits about how much to reveal about why things were penalized and what not, but I think those could be worked around usefully.
Re: My history of (mostly failed) side projects and startups
#67Earlier quoted context omitted.
That was probably me. We have two sides to the webspam team at Google: engineering and manual. We definitely prefer to write algorithms so that we avoid dealing with individual websites--the idea is that you strive to fix the root cause of an issue, not to tackle specific sites. However, if we see a website that violates our guidelines and that gets past the algorithms, we are willing to take manual action. Where pos…
Great to know. Out of curiosity, in this particular case, did you save supposed violations for each site, or did you blacklist all of them based on a few?
I kinda thought one example would make the point. Does it help that much more to give another example? I can look more up. For http://www.bigbadblogdirectory.com/ it looks like you were autogenerating typos not just for websites, but for popular blogs. So http://www.bigbadblogdirectory.com/jeffmatthewsisnotmakingth... looks like it had
(I had to cut out the vast majority of the typos because the comment was too long for HN.)
jeffmatthewsisnotmakingthisup.blogspoot.com, jeffmatthewsisnotmakingthisup.bloyspot.com, jegfmatthewsisnotmakingthisup.blogspot.com, jeffmatthewsisnomakingthisup.blogspot.com, jeffmatthwesisnotmakingthisup.blogspot.com, jeffmatthewsisnotmakingthisup.nlogspot.com, jeffmatthewsisnotmakingthisup.blogspot.ccom, jeffmatthewsisnotmakingthisup.bligspot.com, jeffmatthewsisnotakingthisup.blogspot.com, jeffmatthewsisnotmakinghtisup.blogspot.com, jeffmatthewsisnotmacingthisup.blogspot.com, jdffmatthewsisnotmakingthisup.blogspot.com, jeffmatthewsisnot akingthisup.blogspot.com, ieffmatthewsisnotmakingthisup.blogspot.com, jeffmatthewsisnotmakingthisup/blogspot.com, jeffmatthewsisnotmajingthisup.blogspot.com, jeffmatthewsisnotmakingthishp.blogspot.com, jeff atthewsisnotmakingthisup.blogspot.com, jeffmatthewsisnotmakingthisup.blogspot/com, jeffmatthewwisnotmakingthisup.blogspot.com."
I could post more examples from the other domains, but my point is that this is the sort of thing that users dislike and complain about. If you were a blogger and saw pages like this ranking for your name or your site's name, you probably wouldn't be happy either. From looking at a few domains, I don't think that we overgeneralized from a few pages in this case.
I know that you've moved on and the domains are shut down now. And I'm not trying to be cantankerous. I'm just trying to say that from our point of view there's good reasons to take action on sites like this so that users don't complain to us.
Re: My history of (mostly failed) side projects and startups
#68Earlier quoted context omitted.
Yeah, in retrospect I should have taken it slower and not gotten as close to the line in the first place. It's totally my fault, and I'm not bitter. As you can tell from the OP, I've had a lot of failure, and I similarly learned from this one.
The tricky part is that the math works out something along the lines of there being ~200,000,000 domains and there being ~20,000 Google employees. At a simplistic level that works out to 10,000 domains per Google employee. Which means that even if Google stopped doing everything else and everyone at Google spent all their time talking to webmasters, they'd each have to answer 10,000 peoples' questions about rankings,…
However, I interact with a lot of customers who seem put off by webmaster central. It seems to be a very outdated interface. I understand it's important to be clear and concise when explaining these issues. But if you look around the web 2.0 world at people providing similar information there's a harsh contrast.
Put simply, webmaster central is small text with a dark appearance and little or no graphics. In my experience and testing this harbors a mentality of "This is too complex". Users seem to encounter long wordy pages with no graphics and convinced themselves it's beyond them, before they begin to read.
Making a page lighter and throwing in a few visual aids goes a long way in curbing this issue, as well as making the information easier to understand and more fun to read.
It seems like a small thing, but it scales to become overwhelming when you consider that most people who encounter a page like this and dismiss it at a glance start looking for an email us link or a contact phone number.
This is my experience anyway. Perhaps your results may vary.
Re: My history of (mostly failed) side projects and startups
#69Earlier quoted context omitted.
Great to know. Out of curiosity, in this particular case, did you save supposed violations for each site, or did you blacklist all of them based on a few?
It varies for different cases depending on a lot of factors like severity, impact on users, etc. In the particular case from above, to find out the history of what might have happened, I just picked a domain at random and dug into its history to find the autogenerated pages with tons of typos for each domain. I kinda thought one example would make the point. Does it help that much more to give another example? I can…
Each site took a long time to make actually. They either involved generating a data set from scratch or piecing together and parsing other large data sets. This one in particular, I was crawling the Web for feed discovery and was planning on adding stuff like grouping the best posts by category, etc.
Yeah, would love to know about some others, e.g. japanese2englishdictionary.com, idnscan.com, serverslist.com. Also, did you actually get any complaints about this or was it triggered by some other threshold/thing? On a side note, I still get requests about exposing some of this data, i.e. sites behind ip addresses or lists of domains matching some criteria. In any case, thx for the info!
I can understand the need to take action. I just think it could have been handled better. If typos were the problem, I would have removed them immediately if someone told me, and that could have been automated. In retrospect, it seems pretty obvious, but it wasn't at the time.
Re: My history of (mostly failed) side projects and startups
#70Earlier quoted context omitted.
The tricky part is that the math works out something along the lines of there being ~200,000,000 domains and there being ~20,000 Google employees. At a simplistic level that works out to 10,000 domains per Google employee. Which means that even if Google stopped doing everything else and everyone at Google spent all their time talking to webmasters, they'd each have to answer 10,000 peoples' questions about rankings,…
The videos are great. Also, webmaster central has a ton of great info. However, I interact with a lot of customers who seem put off by webmaster central. It seems to be a very outdated interface. I understand it's important to be clear and concise when explaining these issues. But if you look around the web 2.0 world at people providing similar information there's a harsh contrast. Put simply, webmaster central is sm…
The idea is still percolating, but I think it's got a lot of potential.