Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

141–150 of 365 posts

Re: Only Google is really allowed to crawl the web

#141
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

Are there any actual repercussions for just ignoring robots.txt?

There is if you are doing it for work. For example, your company could get sued if you are found using that data and ignoring the ToS. If you are a public figure, you could get your name tarnished as doing something unethical or the media may call it "hacking". If you are rereleasing the data then you risk getting a takedown notice.

Re: Only Google is really allowed to crawl the web

#143
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

So nobody is going to book air travel? I cant hardly follow what youre even saying besides google=bad.

Re: Only Google is really allowed to crawl the web

#144
post #84

Earlier quoted context omitted.

They don't get individual flight status (what I was talking about) from Sabre or the FAA or Eurocontrol. I didn't get into fares and planned schedules and Google Flights, that's a different topic. I was talking about the big widget you get for queries on status for a particular flight, which is not Google Flights. They have relented in some ways, rolling out stuff in the widget like: "The airline has issued a change…

Oh, status. I was thinking of schedules. Still, what is the point for the consumer of being directed to an airline's terrible status page? And are they even capable of being crawled? Looking at American's site (it was the most ghastly airline that sprang to mind) I don't see how a crawler would be able to deal with it, and indeed the Google snippet for AA flight status, on the aa.com result which is far down in the r…

"what is the point for the consumer of being directed to an airline's terrible status page?"

One example...

If you back up a bit, the widget didn't used to tell you there was a change fee waiver when the flight was full, while aa.com did.

That's an actual, tangible benefit that a consumer might want, worth real money. You can also even often "bid" on a dollar amount to receive if you're willing to change flights. Google doesn't present that info today.

There are more examples. My perspective isn't that Google should lead you to aa.com, but I do feel it's a bit dishonest that the widget is so large it pushes aa.com below the fold. It doesn't need to be that large.

Re: Only Google is really allowed to crawl the web

#145
How about an opt-in search engine cache? One where a domain needs to agree allow their site to be crawled, but as a result also gives said crawler full access? And then that repository would be made publicly available to all search engines to use. Sort of an AP for searches, that would give a base line that wouldn't preclude search engines from going further, but which would certainly lower the cost and network traffic for the search engines and sites that take advantage of it?

Re: Only Google is really allowed to crawl the web

#146
post #136
post #129

Earlier quoted context omitted.

The first link doesn't seem quite conclusive (see the part at the bottom), and also doesn't give evidence that Google's widgets are to blame. The flattening of users could also be due to a general internet-wide reduction in long-form (or even medium-form) non-fiction reading. How are page views for The New York Times? Seems like it should be simple to A/B test, though. Obviously Google could do it themselves by rando…

Edit: Removed bad "simple english graph", thanks. Though the regular english wikipedia traffic is flat from 2016-present. As for NYT, is there a better proxy to compare to? There's no public pageview stats and they have a paywall.

That first graph is Simple English, not English, and is in millions, not billions. They also explicitly call out the methodology change in 2015...

Re: Only Google is really allowed to crawl the web

#147

While I don’t disagree with the idea that all crawlers should have equal access, we also need to address the quality of many crawlers. Google and Microsoft have never hammered any website I’ve run into the ground. Crawlers from other other, smaller, search engines have, to the point where it was easier to just block them entirely. Part of the problem is that sites want search engine to index their site, but not allow…

I think this is a problem which should be solved by automatic rate-limiting and throttling at the application/caching layer (or just individual web server for smaller sites). Requests with a non-browser UA get put into a separate bots-only queue that drains at a rate of ~1/sec or so. If the queue fills up you start sending 429s with random early failures for bots (UA/IP/subnet pairs) that are overrepresented in the t…

I suspect that at least some of the bots use web server response times and response codes as part of the signal for ranking. If your website does not appear capable of handling load then it won't rank as highly, because it is not in their best interests to have search results that don't load.

Re: Only Google is really allowed to crawl the web

#148

Earlier quoted context omitted.

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

Aren't there anti trust laws to prevent this kind of thing?

Antiturst laws are hard to enforce in the United States.

Monopolies themselves aren't illegal. To be convicted of an antitrust violation, a firm needs to both have a monopoly and needs to be using anticompetitive means to maintain that monopoly. The recent "textbook" example was of Microsoft, which in the 90s used its dominant position to charge computer manufacturers for a Windows license for each computer sold, regardless of whether it had Windows installed or was a "bare" PC.

Depending on how you define the market, Google may not even have a monopoly. It's probably dominant enough in web search to count, but if you look at its advertising network it competes with Facebook and other ad networks. In the realm of travel planning (to pick an example from these comments), it's barely a blip.

Furthermore, Google can potentially argue it's not being anticompetitive: all businesses use their existing data to optimize new products, so Google could claim that it not doing so would be an artificial straightjacket.

Re: Only Google is really allowed to crawl the web

#149

Earlier quoted context omitted.

You can screwed any time you book a connecting flight on two different airlines even if the times aren't tight. For instance if one is cancelled. If you use the same airline they will make sure you get to the destination.

That's true, but it can save you a ton of money. You just have to be aware of the risks and plan accordingly. I have typically used this strategy when flying back to the US from the EU. Take an EZJet or similar low cost airline from random small EU city to a larger EU city like Paris, London, Frankfurt, etc... and book the return trip to the US from the larger city. I've also been forced to do this from some EU citie…

Yeah this strategy is good, but you need to allow a long layover like 6 hours if you have to go through immigration and change airports for the connection which happens pretty often with ryanair and ezjet. It’s a big pain, but it does save money.

Re: Only Google is really allowed to crawl the web

#150
post #23

If the shared cache ever became significant enough to matter it would be devastated by marketers, scammers and other abusers. Google employs the groomers that make their index at least tolerable, if still clearly imperfect. Without that cadre of well compensated expertise to win the arms race against such abusers the scheme is not feasible. I suppose this could be crowdsourced if I didn't know about politics and how…

I don't really understand your comment. Marketers, scammers and other abusers already publish to the web with the intention to be included in a crawl. Postprocessing crawl data is already a thing.

Assuming this hypothetical shared crawl cache were to exist, it does not preclude google (and all consumers of that cache) doing their own processing downstream of that cache. Does it?

What's the new attack vector?

Post reply on HN