Live data from Hacker News

Suddenly, Hacker News is not the first result for 'Hacker News'

google.com

101–110 of 216 posts

Re: Suddenly, Hacker News is not the first result for 'Hacker News'

#101
post #63
post #29

Does anyone else find it disturbing that google employees are bending over for pg/hn? Seriously, if any other webmaster blocked googles bots they wouldn't change their algorithms to accommodate, or see how they could use less of our resources. Its pg's fault not googles, and I dont see why they should care. Maybe from their standpoint it would be more beneficial to google users who are used to typing in 'hacker news'…

>Does anyone else find it disturbing that google employees are bending over for pg/hn? Umm...no. They're not tweaking the algo, they're explaining to PG that he should stop blocking the crawlers, or that he should verify the site in google webmaster tools, then change the crawl rate.

Umm...Yes.

Seems like "bending over" to me:

"but I'm pinging the right people to ask whether we can get this fixed pretty quickly"

Re: Suddenly, Hacker News is not the first result for 'Hacker News'

#102
post #25

Earlier quoted context omitted.

I would like to set that crawl rate but do not see why I must register at Google to do so. Why can't Google support the Crawl-Delay directive in robots.txt for this?

My plane is about to take off, but very briefly: people sometimes shoot themselves in the foot and get it way, way wrong. Like "crawl a single page from my website every five years" wrong. Crawl-Delay is (in my opinion) not the best measure. We tend to talk about "hostload," which is the inverse: the number of simultaneous connections that are allowed.

I do understand the thought but I think it is not a good gesture to do. You could always cap crawl-delay at a reasonable maximum and additionally allow people to fix mistakes through the webmaster tools (eg if they told your bots to stay away for a long time but in the meantime want to revert that).

Maybe instead that hostload could be parsed from robots.txt? It sure seems like the better mechanic to tweak for load issues (while traffic/bandwidth issues are still unresolved).

Re: Suddenly, Hacker News is not the first result for 'Hacker News'

#103
post #99
post #47

Earlier quoted context omitted.

Seeing threads like this remind me that HN is still a pretty tight-knit community of real people doing real things. It's good to see this stuff sometimes. Thanks, Matt! edit: And then reading some of the other threads on this topic is a bit...something. Guys, can you calm the conspiracy theory nonsense a bit? Please? If you're not on this site very much, you might not realize that Matt pops into almost every thread w…

"This isn't HN getting some sort of preferential treatment, this is just the effect of having a userbase full of hackers" Of course it's preferential treatment. And if you scan the last month or two of Matt's comments they are general in nature and not specific as in: "I think I know what the problem is; we're detecting HN as a dead page. It's unclear whether this happened on the HN side or on Google's side, but I'm…

He's done the same thing for nearly everyone that's asked about something Google-related here. So yes, everyone on HN gets special treatment.

Answering people's individual questions doesn't scale to the entire Internet, so Google really has no choice but to address problems on a case-by-case basis. In this case, Matt reads HN and personally wants to solve the problem. That's the only way Google could possibly work, so that's how they do it.

Re: Suddenly, Hacker News is not the first result for 'Hacker News'

#104

Earlier quoted context omitted.

My plane is about to take off, but very briefly: people sometimes shoot themselves in the foot and get it way, way wrong. Like "crawl a single page from my website every five years" wrong. Crawl-Delay is (in my opinion) not the best measure. We tend to talk about "hostload," which is the inverse: the number of simultaneous connections that are allowed.

I would think that the number of people who (a) know how to create a valid robot.txt file, (b) have some idea of how to use the "crawl-delay" directive and (c) write a "shoot-themselves-in-the-foot" worthy error is vanishingly small.

Considering it's Google and we're talking about almost the whole population of the Earth, the vanishingly small percentage of the entire population of the planet would still be at least hundreds of thousands of people.

Re: Suddenly, Hacker News is not the first result for 'Hacker News'

#105
post #25

Earlier quoted context omitted.

I would like to set that crawl rate but do not see why I must register at Google to do so. Why can't Google support the Crawl-Delay directive in robots.txt for this?

My plane is about to take off, but very briefly: people sometimes shoot themselves in the foot and get it way, way wrong. Like "crawl a single page from my website every five years" wrong. Crawl-Delay is (in my opinion) not the best measure. We tend to talk about "hostload," which is the inverse: the number of simultaneous connections that are allowed.

Matt, how does a sitemap fit into this? If I'm not mistaken, you can suggest some refresh rate there, too. Do you take that into account?

Re: Suddenly, Hacker News is not the first result for 'Hacker News'

#106

Earlier quoted context omitted.

My plane is about to take off, but very briefly: people sometimes shoot themselves in the foot and get it way, way wrong. Like "crawl a single page from my website every five years" wrong. Crawl-Delay is (in my opinion) not the best measure. We tend to talk about "hostload," which is the inverse: the number of simultaneous connections that are allowed.

I would think that the number of people who (a) know how to create a valid robot.txt file, (b) have some idea of how to use the "crawl-delay" directive and (c) write a "shoot-themselves-in-the-foot" worthy error is vanishingly small.

I alluded to some of the ways that I've seen people shoot themselves in the foot in a blog post a few years ago: http://www.mattcutts.com/blog/the-web-is-a-fuzz-test-patch-y...

"You would not believe the sort of weird, random, ill-formed stuff that some people put up on the web: everything from tables nested to infinity and beyond, to web documents with a filetype of exe, to executables returned as text documents. In a 1996 paper titled "An Investigation of Documents from the World Wide Web," Inktomi Eric Brewer and colleagues discovered that over 40% of web pages had at least one syntax error".

We can often figure out the intent of the site owner, but mistakes do happen.

Re: Suddenly, Hacker News is not the first result for 'Hacker News'

#107
post #95

Earlier quoted context omitted.

I would think that the number of people who (a) know how to create a valid robot.txt file, (b) have some idea of how to use the "crawl-delay" directive and (c) write a "shoot-themselves-in-the-foot" worthy error is vanishingly small.

As opposed to --for illustrative purposes-- the vanishingly small number of people who know how to block IP addresses and manage to get their site to disappear from Google's listings? ::cough:: A few years ago, I did pretty much the same thing myself. Thankfully the late summer was our slow season and the site recovered pretty quickly from my bone-headed move, but the split second after I realized what I've done was…

Perhaps I'm making an error assuming that a website as influential as Hacker News has a "real live Webmaster" to do things like write robot.txt files.

Re: Suddenly, Hacker News is not the first result for 'Hacker News'

#108

Earlier quoted context omitted.

Yes, it's VERY disturbing, no question about it. Matt can sound all helpful, and cuddly, and oxytocin-inducing all he wants, but he's basically stepping in and helping 1 site over another, when the whole process should NOT favor anyone. There are lots of webmasters out there that do stupid stuff like override a robots.txt accidentally, yet I don't see Matt sending them an email, asking them to check up on it.. Come'n…

Are you kidding? Matt is basically a superhero who helps anyone he can find. He's solved problems for people complaining on Twitter, on Google+, he hosts regular video office hours when anyone can come to him with problems, etc. It may be the first time YOU'VE seen him rush in to help someone, but this is Matt Cutts' standard MO.

No, he's not a super hero.. and no he doesn't help everyone that asks for it. I've brought up many issues over the years to him via Twitter, blog comments, emails, etc, and none of them.. count them.. ZERO have been addressed by him. He just gives preferential treatment to those sites that make a lot of noise, and might give Google or Google's quality team a lot of negative press. That's all he does. I've even brought up 100% apparent, clear spammy sites and he doesn't respond, and doesn't do anything about them. They're still there.

Lots of folks here like to give Matt Cutts this aura of an angel, or some sort of saint. Those who have known his actions over the past 5-10 years know better. Just ask guys like Aaron Wall and Rand. Matt Cutts is just a pawn that is there to do damage control for Google. That's all he is.

Re: Suddenly, Hacker News is not the first result for 'Hacker News'

#109
post #15

Earlier quoted context omitted.

I sent you an email about this. (A couple weeks ago I banned all Google crawler IPs except one. Crawlers are disproportionately bad for HN's performance because HN is optimized to serve recent stuff, which is usually in memory.)

Hi Paul, A site can be crawled from any number of Googlebot IP addresses, and so blocking all except one doesn't help in throttling crawling. If you verify the site in Webmaster Tools, we have a tool you can use to set a slower crawl rate for Googlebot, regardless of which specific IP address ends up crawling the site. Let me know if you need more help. Edit Detailed instructions to set a custom crawl rate: 1. Verify…

Could google vary the crawling rate on each site and see what effects that has on response times, and develop an algorithm to adjust crawl speed so as not to affect site performance too much? If google starts crawling a site and notices sequential crawl requests are answered in .5s w/ .1s stddev and it starts crawling with 10 parallel connections and the answers are 2s w/ 1s stddev, clearly that's a problem because user experience for real people will be impacted. Maybe google could automatically email webmaster@ and notify them of performance issues it sees when crawling.

Another thing that might help google is for them to announce and support some meta tag that would allow site owners (or web app devs) to declare how likely a page is to change in the future. Google could store that with the page metadata and when crawling a site for updates, particularly when rate limited via webmaster tools, it could first crawl those pages most likely to have changed. Forum/discussion sites could add the meta tags to older threads (particularly once they're no longer open for comments) announcing to google that those thread pages are unlikely to change in the future. For sites with lots of old threads (or lots of pages generated from data stored in a DB and not all of which can be cached), that sort of feature would help the site during google crawls and would help google keep more recent pages up to date without crawling entire sites.

Post reply on HN