Live data from Hacker News

Applebot, the web crawler for Apple

support.apple.com

91–94 of 94 posts

Re: Applebot, the web crawler for Apple

#91
post #52

>If robots instructions don't mention Applebot but do mention Googlebot, the Apple robot will follow Googlebot instructions. So if I set in my robots.txt to disallow all bots except Googlebot, Applebot will index anyway? I don't think I like that precedent.

I've said it before, and I'll say it again: blocking all but certain bots is the best way to block innovation. It's bad for you, it's bad for the ecosystem. Bad for the ecosystem because incumbents that want to respect robots.txt to propose new services will have an unfair disadvantage. It's bad for you because it'll just give more power to Google et al regarding your incoming traffic (and you'll have to follow their…

But that is a choice that a website operator has the authority to make given current standards. Apple is ignoring the wishes of a website operator and piggybacking on the trust that a website operator has place in GoogleBot.

Re: Applebot, the web crawler for Apple

#92
post #88
post #52

Earlier quoted context omitted.

I've said it before, and I'll say it again: blocking all but certain bots is the best way to block innovation. It's bad for you, it's bad for the ecosystem. Bad for the ecosystem because incumbents that want to respect robots.txt to propose new services will have an unfair disadvantage. It's bad for you because it'll just give more power to Google et al regarding your incoming traffic (and you'll have to follow their…

Note that serving bots, especially for media-rich sites, eats heavily into a finite resource. For those of us whose resources are especially tight, blocking Yandex, Baidu and MSN may be very helpful, your ideals notwithstanding.

Putting anything on the internet exposes your "finite resources" to attackers. I'm sorry, but robots.txt is just a courtesy offered to you, and if you're doing serious business, you shouldn't rely on it. Use authentication, rate-limiting, captchas, DDoS-protecting CDNs, and enforce the limitations at the source, don't rely on nice people respecting your robots.txt.

Re: Applebot, the web crawler for Apple

#93

Earlier quoted context omitted.

That's what you get when many people code with only one company in mind. It happened in the past and it is happening again, whether we like it or not. Even Microsoft was kinda masquerading early versions of Edge as Chrome.

"Even Microsoft was kinda masquerading early versions of Edge as Chrome" I would assume that was more to keep the tech press from seeing it.

I don't know about the not public betas, but as soon as there was a public Windows 10 TP and you could swap IE's engine to Edge I'm pretty sure they had "Edge" at the end of the user-agent, but they had the rest of the Chrome UA so most of the websites out there considered it as being Chrome.

I may also be completely wrong :)

Re: Applebot, the web crawler for Apple

#94
post #89

Earlier quoted context omitted.

I wonder what are the legalities of this type of discrimination? Retail businesses aren't allowed to arbitrarily refuse service, they have to follow certain rules and be consistent.

Is there legal backing to enforce obeying robots.txt, or is it a guideline?

There are some court cases that involve robots.txt but I'm not sure because there was a lot of reading to do and I was hoping someone who knew would just come along and summarize it for us :)
Post reply on HN