Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

291–300 of 365 posts

Re: Only Google is really allowed to crawl the web

#291
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

> Google lets the data dictate which markets to enter and on one hand they can jack up advertising fees on customers/competitors and unfairly build their own service into search above both ads and organic results.

Just like Amazon with Amazon Basics.

Re: Only Google is really allowed to crawl the web

#292
post #237

Earlier quoted context omitted.

consumers are in this case the advertisers. google has a monopoly on search ads and does enforce it, being a drain on the economy since in many fields you only succeed if you spend on search ads

googles answer to this at yesterdays hearing.. Search isnt a single category. If you break it down, they arent a monopoly. For example. 1/2 of PRODUCT SEARCHES begin on Amazon. It's probably hard to argue Google as a monopoly if who they see as their main competitor has half the market share.

That's so disingenuous there should be a new term for it.

Re: Only Google is really allowed to crawl the web

#293

Earlier quoted context omitted.

Broadly speaking, robots.txt files are often ignored. I used to run a fairly large job ad scraping organization, and we would be hired by companies (700 of the fortune 1000 used us) to scrape the job ads from their career pages, and then post those jobs on job boards. 99 of 100 times, the robots file would disallow us to scrape. Since we were being paid by that company's HR team to scrape, we just ignored it because…

> Broadly speaking, robots.txt files are often ignored. If you wanna go nuclear on people who do that, include an invisible link in your html and forbid access to that URL in your robots.txt, then block every IP who accesses that URL for X amount of time. Don't do this if you actually rely on search engine traffic though. Google may get pissed and send you lots of angry mail like "There's a problem with your site".

We would occasionally have customer try doing that. AWS has lots of IP addresses :-).

Re: Only Google is really allowed to crawl the web

#294
post #195

Earlier quoted context omitted.

How often were you searching?

I was regularly searching, but I was rarely indexing any of the sites. I struggled to even get an initial index of many of the sites, due to being blocked or being reported.

If you were getting blocked on an initial load, you were either hit with rate limiting or an unrecognized user agent

Re: Only Google is really allowed to crawl the web

#295

Earlier quoted context omitted.

That's true, but it can save you a ton of money. You just have to be aware of the risks and plan accordingly. I have typically used this strategy when flying back to the US from the EU. Take an EZJet or similar low cost airline from random small EU city to a larger EU city like Paris, London, Frankfurt, etc... and book the return trip to the US from the larger city. I've also been forced to do this from some EU citie…

The difference is mind-boggling in some cases. On one trip in 2019 I had the following coach fair choices for SFO - Moscow return trip tickets booked 3 weeks prior to departure. * UA or Lufthansa round trip (single carrier) $3K * UA round trip SFO - Paris + Aeroflot round trip Paris - Moscow: $1K No amount of search could reduce the gap. I went with the second option. The gap is even bigger if you have a route with m…

https://www.airtreks.com/ will do this for you with a person. phenomenal service.

To anyone from airtreks, I love you so much!

Re: Only Google is really allowed to crawl the web

#296

> Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. This was eyebrow-raising. Actually seems like this is not (any longer?) true: https://census.gov/robots.txt : User-agent…

Actually, thinking more about this, I think they might be misconfigured, because they clearly don't want robots touching /cgi-bin etc (reasonable!) but they are actually only asking the named robots to do that, all other bots have no guidance about what not to touch

Re: Only Google is really allowed to crawl the web

#297

Earlier quoted context omitted.

I usually recommend setting only Google/Bing/Yandex/Baidu etc to Allow and everything else to Disallow. Yes, the bad bots don't give a fuck, but even the non-malicious bots (ahrefs, moz, some university's search engine etc) don't bring any value to me as a site owner, take up band width and resources and fill up logs. If you can remove them with three lines in your robots.txt, that's less noise. Especially universiti…

> they're hammering your server Why can't you just ratelimit IPs that are "too active" for your server to handle?

CGNAT

Re: Only Google is really allowed to crawl the web

#298
post #98

Earlier quoted context omitted.

Aren't there anti trust laws to prevent this kind of thing?

Anti-trust in the US tend to not hit the big tech players as much they do other sectors. Also there is actually a debate in the judicial system about the extent of Anti trust laws themselves.

Chicago school basically published a bunch of position papers that made feudal corporations a legal entity that "aren't monopolies" because the give things away for free. Because the consumer isn't paying, it can't be bad.

Re: Only Google is really allowed to crawl the web

#299

Earlier quoted context omitted.

Isn't that the website owners right though? I'm not sure I understand the problem here. If Google is taking traffic and reducing revenue, a company can deny in robots.txt. Google will actually follow those rules - unlike most others that are supposedly in this 2nd class.

> Isn't that the website owners right though? No. The internet is public. Publishers shouldn't get any say in who accesses their content or how they do it. As far as I'm concerned, the fact that they do is a bug.

No, it's not. I can setup a login page and keep you out if I want. And I can do it however I want.

Re: Only Google is really allowed to crawl the web

#300
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

I think the flight arrivals/departures is a bad example. A good example might be putting flights.google.com on the first page or even allow it to exist.
Post reply on HN