Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

121–130 of 365 posts

Re: Only Google is really allowed to crawl the web

#121
post #17

I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…

It's hilarious to think there exists people who think googlebot does not get special treatment from website operators. Here is an experiment you can do in a jiffy, write a script that crawls any major website and see how many URL fetches it takes before your IP gets blocked. Googlebot has a range of IP addresses that it publicly announces so websites can whitelist them.

I have never had that problem running screaming frog on big brand sites apart from one or two times.

Re: Only Google is really allowed to crawl the web

#122
post #105
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

Is it legal for a government entity to issue a robots.txt like that? Maybe the line between use and abuse hasn't been delinated as well as it needs to be.

> Is it legal for a government entity to issue a robots.txt like that?

I may be wrong (this isn't my area), but I was under the impression that robots.txt was just an unofficial convention? I'm not saying people should ignore robots.txt, but are there legal ramifications if ignored? I'm not asking about techniques sites use to discourage crawlers/scrapers, I'm specifically wondering if robots.txt has any legal weight.

Re: Only Google is really allowed to crawl the web

#123
Can we take a moment to talk about this club's business model?

There's not even any information to see what the "private forum access" that you have to pay for is about, what kind of people are in it...or even to know about what happens with the money.

For me, this sounds like a scam.

I mean, no information about any company. No imprint. No privacy policy. No non-profit organization. And just a copy/paste wordpress instance.

I mean, srsly. I am building a peer-to-peer network that tries to liberate the power of google, specifically, and I would not even consider joining this club. And I am the best case scenario of the proposed market fit.

Re: Only Google is really allowed to crawl the web

#125
post #50
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

>However, that sort of search used to (most times) lead to a visit to the airline web site. I don't think that's correct. In the old days you'd either call a travel agent or use an aggregator like expedia. Google muscles out intermediaries like Expedia, Yelp, and so on. It's not likely much better or worse for the end user or supplier. Just swapping one middleman for another.

I don't think google muscling out intermediaries like Expedia is a good thing.

Just for example, Expedia is probably 5% of Google's total revenue and Google doesn't like slim margin services by and large that can't be automated.

Travel is fairly high-touch - people centric. It doesn't fit Google's "MO".

But... its shitty that google can play all sides of the markets while holding people ransom to mass sums of money to pay to play on PPC where google doesn't... i think that's where the problem shines.

In essence, you're advocating that eBay goes away because google could do it... they could.. and eBay is technically just an intermediary, but do we want everything to be googlefied?

Google bought up/destroyed other aggregators - remember the days of fatwallet, priceline, pricewatch, shopzilla and such when they used to focus on discounts/coupons/deals and now they're moving more towards rewards/shopping/experience - it used to be i could do PPC on pricewatch and reach millions of shoppers are a reasonable rate, but now that google destroyed them all, the PPC rate on "goods" is absurdly high and not having an affordable market means only the amazons and walmarts can really afford to play...

it used to be you could niche out, but even then, that's getting harder

Re: Only Google is really allowed to crawl the web

#126

I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…

When you operate commercial sites at scale, bots are a real thing you spend real engineering hours thinking about and troubleshooting and coding to solve for.

And yes, that means google gets special treatment.

Think about the model for a site like stackoverflow. The longest of long tail questions on that site: what’s the actual lifecycle of that question?

- posted by a random user - scraped by google, bing, et al - visited by someone who clicked on a search result on google - eventually, answered - hopefully, reindexed by google, bing et al - maybe never visited again because the answer now shows up on the google SERP

In the lifetime of that question how many times is it accessed by a human, compared to the number of times it’s requested and rerequested by an indexing bot?

What would be the impact on your site of three more bots as persistent as google bot? Why should you bother with their requests?

So yes, sites care about bot traffic and they care about google in particular.

Re: Only Google is really allowed to crawl the web

#127
post #105
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

Is it legal for a government entity to issue a robots.txt like that? Maybe the line between use and abuse hasn't been delinated as well as it needs to be.

Is failure to honor a robots.txt a crime? Or rather, would it be unlawful to spoof a user agent to access this publicly available data? After the linkedin [0] case it seems reasonable to think not.

[0]: https://www.eff.org/deeplinks/2019/09/victory-ruling-hiq-v-l...

Re: Only Google is really allowed to crawl the web

#128
post #84
post #70

Earlier quoted context omitted.

You're wrong on a lot of facts here. Google Flights doesn't get its data just by crawling, they get it from Sabre, the FAA, Eurocontrol, etc. Airlines are, obviously, extremely pleased to disseminate this information. Google Flights "gives back" in the exact same way as any other travel outlet: they book passengers. As for Wikipedia, the WMF is quite happy that most of their traffic is now served by Google. WMF is in…

They don't get individual flight status (what I was talking about) from Sabre or the FAA or Eurocontrol. I didn't get into fares and planned schedules and Google Flights, that's a different topic. I was talking about the big widget you get for queries on status for a particular flight, which is not Google Flights. They have relented in some ways, rolling out stuff in the widget like: "The airline has issued a change…

Oh, status. I was thinking of schedules. Still, what is the point for the consumer of being directed to an airline's terrible status page? And are they even capable of being crawled? Looking at American's site (it was the most ghastly airline that sprang to mind) I don't see how a crawler would be able to deal with it, and indeed the Google snippet for AA flight status, on the aa.com result which is far down in the results page, just says "aa.com uses cookies" which is about what you'd expect.

In this case, I want to be sent literally anywhere but aa.com.

Re: Only Google is really allowed to crawl the web

#129
post #108

Earlier quoted context omitted.

> You see this recently with Wikipedia. Google's widgets have been reducing traffic to Wikipedia pretty dramatically. Wikipedia visitors, edits, and revenue are all increasing, and the rate that they're increasing is increasing, at least in the last few years. Is this a claim about the third derivative? > Enough so that Wikipedia is now pushing back with a product that the Googles of the world will have to pay for. T…

See https://searchengineland.com/wikipedia-confirms-they-are-ste... from 2015. Google's widgets that present Wikipedia data do reduce visitors to Wikipedia. Or see page views on English Wikipedia from 2016-current: https://stats.wikimedia.org/#/en.wikipedia.org/reading/total... Looks pretty flat, right? Does that seem normal? As for Wikimedia Enterprise, you do have to read between the lines a bit. "The focus is on o…

The first link doesn't seem quite conclusive (see the part at the bottom), and also doesn't give evidence that Google's widgets are to blame.

The flattening of users could also be due to a general internet-wide reduction in long-form (or even medium-form) non-fiction reading. How are page views for The New York Times?

Seems like it should be simple to A/B test, though. Obviously Google could do it themselves by randomly taking away the widget, but would could also see whether referrals from non-Google search engines (though they are themselves a tiny percentage) continue to increase while Google remains flat.

Re: Only Google is really allowed to crawl the web

#130

Earlier quoted context omitted.

Be careful when using Google Flight, last time I checked they use significantly less margins between flights so it’s shorter trips but much riskier.

You can screwed any time you book a connecting flight on two different airlines even if the times aren't tight. For instance if one is cancelled. If you use the same airline they will make sure you get to the destination.

That's true, but it can save you a ton of money. You just have to be aware of the risks and plan accordingly.

I have typically used this strategy when flying back to the US from the EU. Take an EZJet or similar low cost airline from random small EU city to a larger EU city like Paris, London, Frankfurt, etc... and book the return trip to the US from the larger city. I've also been forced to do this from some EU cities since there was no connecting partner with a US airline.

Post reply on HN