Live data from Hacker News

Applebot, the web crawler for Apple

support.apple.com

81–90 of 94 posts

Re: Applebot, the web crawler for Apple

#81
post #27

This is interesting. A while back, I think either Cook or Jobs mentioned that Apples makes PRODUCTS and doesn't sell ADS. If that's true (and stays true) AND this is the beginning of a search engine for them, it's going to be VERY interesting to see what it looks like.

I'm assuming you are remembering Cook's open letter on privacy http://www.apple.com/privacy/ "Our business model is very straightforward: We sell great products. We don’t build a profile based on your email content or web browsing habits to sell to advertisers." This obviously is not saying Apple doesn't sell ads. It's saying their core business model is selling products. The users are not the product.

It does not say they don't collect data thought. And I'm pretty sure not even Google sells data to advertisers, they sell ad space. It's just a PR stunt.

Re: Applebot, the web crawler for Apple

#82

Earlier quoted context omitted.

Apple still collects data on you. Google knows what sites you browse on the Web and what Web sites you search for. Apple knows what apps you search for in the App Store and what you buy in the store, and probably usage statistics. At some point, if not already, that data will be used to profile you and generate revenue maximizing suggestions of goods and services to buy in the App Store.

You're completely missing the point. Google makes 90% of their revenue from advertising. Apple makes <5% of their revenue from App Store sales.

So you're ok with your data being collected, as long as they don't make money of it?

Re: Applebot, the web crawler for Apple

#83

Apple trying to get do what Google does faster, than Google doing what Apple does. In other words: Both Apple and Google in the mobile OS and device space. Both Apple and Google in the (mobile) search space.

Has anyone really been far even as decided to use even go want to do look more like?

Re: Applebot, the web crawler for Apple

#84

>If robots instructions don't mention Applebot but do mention Googlebot, the Apple robot will follow Googlebot instructions. So if I set in my robots.txt to disallow all bots except Googlebot, Applebot will index anyway? I don't think I like that precedent.

Why? Applebot is attempting to adhere to the spirit of robots.txt rather than the letter. If the site owner cares otherwise, don't just explicitly name one bot. Is `user-agent: *bot` valid?

Re: Applebot, the web crawler for Apple

#85
post #52

>If robots instructions don't mention Applebot but do mention Googlebot, the Apple robot will follow Googlebot instructions. So if I set in my robots.txt to disallow all bots except Googlebot, Applebot will index anyway? I don't think I like that precedent.

I've said it before, and I'll say it again: blocking all but certain bots is the best way to block innovation. It's bad for you, it's bad for the ecosystem. Bad for the ecosystem because incumbents that want to respect robots.txt to propose new services will have an unfair disadvantage. It's bad for you because it'll just give more power to Google et al regarding your incoming traffic (and you'll have to follow their…

I wonder what are the legalities of this type of discrimination? Retail businesses aren't allowed to arbitrarily refuse service, they have to follow certain rules and be consistent.

Re: Applebot, the web crawler for Apple

#86
post #21

Earlier quoted context omitted.

I'm not an expert but my guess is: limiting bot traffic, but keeping the site available for the most popular search engine.

My experience is that the worst bots don't respect robots.txt anyway. Getting crawled by the major search engines typically isn't that bad, they tend to know what they're doing. Getting hammered by some crappy local search engine is what's annoying. We don't limit any bots, except once where we completely blocked Eniro in our firewall. Google, Bing and a ton of other could index at the same time, with no issue. Eniro…

I thought FB was the internet. Googlebot is just the Kleenex of indexers.

Re: Applebot, the web crawler for Apple

#87
post #65

Earlier quoted context omitted.

Hey, I don't think it's appropriate to post publicly information that somebody divulged to you in confidence.

I don't think email is considered confidential by default. Had this recruiter made any confidentiality request, I would have tried my best to honor it. Instead, he seemed more interested in spreading the word that they were entering the search game. Also, Apple has not been exactly hiding their growing interest in search. They rarely let their engineers speak in public, but they were on stage this year at Lucene Revo…

It's not clear-cut, that's for sure. I just wanted to post a different perspective.

Thing is, if you asked the sender for permission to post pieces of their email, they'd probably say no. It seems a bit gauche to say posting is okay because "nobody told me not to."

Re: Applebot, the web crawler for Apple

#88
post #52

>If robots instructions don't mention Applebot but do mention Googlebot, the Apple robot will follow Googlebot instructions. So if I set in my robots.txt to disallow all bots except Googlebot, Applebot will index anyway? I don't think I like that precedent.

I've said it before, and I'll say it again: blocking all but certain bots is the best way to block innovation. It's bad for you, it's bad for the ecosystem. Bad for the ecosystem because incumbents that want to respect robots.txt to propose new services will have an unfair disadvantage. It's bad for you because it'll just give more power to Google et al regarding your incoming traffic (and you'll have to follow their…

Note that serving bots, especially for media-rich sites, eats heavily into a finite resource.

For those of us whose resources are especially tight, blocking Yandex, Baidu and MSN may be very helpful, your ideals notwithstanding.

Re: Applebot, the web crawler for Apple

#89
post #52

Earlier quoted context omitted.

I've said it before, and I'll say it again: blocking all but certain bots is the best way to block innovation. It's bad for you, it's bad for the ecosystem. Bad for the ecosystem because incumbents that want to respect robots.txt to propose new services will have an unfair disadvantage. It's bad for you because it'll just give more power to Google et al regarding your incoming traffic (and you'll have to follow their…

I wonder what are the legalities of this type of discrimination? Retail businesses aren't allowed to arbitrarily refuse service, they have to follow certain rules and be consistent.

Is there legal backing to enforce obeying robots.txt, or is it a guideline?

Re: Applebot, the web crawler for Apple

#90
post #87

Earlier quoted context omitted.

I don't think email is considered confidential by default. Had this recruiter made any confidentiality request, I would have tried my best to honor it. Instead, he seemed more interested in spreading the word that they were entering the search game. Also, Apple has not been exactly hiding their growing interest in search. They rarely let their engineers speak in public, but they were on stage this year at Lucene Revo…

It's not clear-cut, that's for sure. I just wanted to post a different perspective. Thing is, if you asked the sender for permission to post pieces of their email, they'd probably say no. It seems a bit gauche to say posting is okay because "nobody told me not to."

It's a good point and since journalists are already asking about it, I now wish I hadn't even posted it. I think Apple is not trying to hide the fact that they are looking for search and machine learning people, but the press will surely get it wrong trying to triangulate a vague one-liner from a recruiting email.
Post reply on HN