Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

311–320 of 365 posts

Re: Only Google is really allowed to crawl the web

#311

On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…

On the other hand, I do not want my site to go down thanks to a few bad 'crawlers' that fork() a thousand http requests every second and take down my site, forcing me to do manual blocking or pay for a bigger server/scale-out my infrastructure. Why should I have to serve them?

Agree... this entire argument seems to think I give a rats ass about 9000 different crawlers that give me literally zero benefit and only waste server resources. Most of those crawlers are for ad-soaked piss poor search engines. I'd rather just block them all and allow crawlers that don't know how to behave.

Re: Only Google is really allowed to crawl the web

#312
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

> The airline web site could then present things Google can't do. Like "hey, we see you haven't checked in yet" or "TSA wait times are longer than usual" or "We have a more-legroom seat upgrade if you want it". If I'm a passenger, there's plenty of ways for airlines to notify me. If I'm searching for a flight status online, it's because I'm picking someone up. If I want more information, I'll click through. I don't s…

It's a progression. It's not a huge problem, by itself, for either. But Google shareholders want to continue the same YoY gains. The only cash cow is search, so they continue to take screen real estate that used to go to others, and take it for themselves. Whether that's more ads, or more widgets, or whatever.

Yes, it's legal. But it does reduce visitor interactions for those sites. Reduced visitor interactions isn't good for web sites...it takes away incentives, reduces brand value, reduces revenue. Eventually, that is not great for consumers.

Ever been a middleman? Squeeze your suppliers enough, and you kill them. The next supplier will pre-emptively cut quality, features, etc, because they know you're going to try and squeeze them to death.

"If I'm searching for a flight status online, it's because I'm picking someone up"

That's one use case, it's not all of them. There's middle ground too, like the bones they throw Wikipedia in the form of links for more info.

Re: Only Google is really allowed to crawl the web

#313
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

isn't the solution to disaffective aqui - hiring the restoration of the ability to IPO companies like ITA Software so there's another way to reach financial security for talented programmer - entrepreneurs? for that matter, do the conditions for recreating the frequency of lower level programmer millionaires (hardly a family home debt free today) like Microsoft created, require the recreation also of senior executive abuse of option schemes? It seems to me that making it a reasonable chance of becoming at least financially secure for not irrational amounts of dedication and 90 weeks, and underwriting that with the greater robustness of larger companies and hence livable salaries, instead of trying to sustain the apparent startup free for all figuring to a common heat death of the advertising budgetary universe?

Re: Only Google is really allowed to crawl the web

#314

On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…

I wonder what happens to RSS feeds in this situation. Programs I run that process RSS feeds will just fetch them over HTTP completely headlessly, so if there are any CAPTCHAs, I'm not going to see them.

In my experience, those either get detected(?) and let through (rss can be agressively cached after all) or you're out of luck and the website owner set up e.g. wordpress (which automatically included rss URLs) but did not configure cloudflare to let rss through.

Re: Only Google is really allowed to crawl the web

#315

Earlier quoted context omitted.

That's true, but it can save you a ton of money. You just have to be aware of the risks and plan accordingly. I have typically used this strategy when flying back to the US from the EU. Take an EZJet or similar low cost airline from random small EU city to a larger EU city like Paris, London, Frankfurt, etc... and book the return trip to the US from the larger city. I've also been forced to do this from some EU citie…

Yeah this strategy is good, but you need to allow a long layover like 6 hours if you have to go through immigration and change airports for the connection which happens pretty often with ryanair and ezjet. It’s a big pain, but it does save money.

in my ideal world the software ITA wrote for airlines and is now owned by Google would be in the hands of consumers and the airlines could have adapted to shifts in demand probably without the need for abrupt cessation of services and human fatigue on industry employees caused when route optimisation analysis tempts executives with what I suspect are ultimately fictitious net present savings.

Re: Only Google is really allowed to crawl the web

#316

Earlier quoted context omitted.

Are there any actual repercussions for just ignoring robots.txt?

There is if you are doing it for work. For example, your company could get sued if you are found using that data and ignoring the ToS. If you are a public figure, you could get your name tarnished as doing something unethical or the media may call it "hacking". If you are rereleasing the data then you risk getting a takedown notice.

robots.txt is not a terms of service. Even if it was, it wouldn't be enforceable for a public website. You would need to prove that a web crawler is maliciously causing disruption to your service, and that is not easy.

Re: Only Google is really allowed to crawl the web

#319

Earlier quoted context omitted.

Even before it gets to that point, they routinely display snippets off regular websites and show ads next to it. Keeping users from clicking through to organic results helps them generate more revenue.

Here’a a thought: most companies don’t actually want to serve a ton of extra pages. For example, airlines just want to fly passengers. They don’t care who puts those butts in seats and they would fully acknowledge that they aren’t able to deliver a better flight search than Google can. I mean, sure, some small team of web developers at every airline is pissed, but the CEO needs butts in seats to keep the pilot and se…

Sure, but if you aren’t controlling the experience of getting butts in seats, it’s harder to upsell and make even _more_ money from those butts.

Re: Only Google is really allowed to crawl the web

#320
post #237

Earlier quoted context omitted.

googles answer to this at yesterdays hearing.. Search isnt a single category. If you break it down, they arent a monopoly. For example. 1/2 of PRODUCT SEARCHES begin on Amazon. It's probably hard to argue Google as a monopoly if who they see as their main competitor has half the market share.

That's so disingenuous there should be a new term for it.

gSplain or gWash
Post reply on HN