Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

321–330 of 365 posts

Re: Only Google is really allowed to crawl the web

#321
I believe there is a lot of hidden fight behind the scenes for Google to monopolize the web.

There are a lot of expectations from the public that Google maintains. Apart from delivering the best search results, they have to respect robots.txt, limit crawling frequency, deliver search results really fast, etc.

At the same time, lots of people want to game Google and spam search results so they can make money with ads. The competition is not fair at all - the websites can tell which traffic came from GoogleBot and craft legitimate responses while Google is not allowed to publicly crawl the website with a fake User Agent.

Many websites are concerned about their data being crawled (like Amazon, they certainly wouldn't be happy if all their price information is dumped as a database), and Google has to make sure that no robot can crawl too much of a website by automatically searching Google, and to that end they invented reCATPCHA.

It's not easy at all to build a search engine that behaves responsibly both to websites and to users. It's the ability to deal with all these matters really well that gives Google the monopoly power.

Re: Only Google is really allowed to crawl the web

#322
post #301

Earlier quoted context omitted.

Unlikely that Wikipedia will be able to charge for content, seeing as all of their content is CC-BY-SA licensed. https://en.wikipedia.org/wiki/Wikipedia:Licensing_update They may be able to charge for bandwidth (if you want to use a Wikipedia image, you can use Wikipedia's enterprise CDN instead of their own), but their licensing allows me to rehost content as long as I follow the attribution & sublicensing terms. Go…

Fine, the content is free. But if your crawlers want access to the content, then pay! Simple as that.

Will it be a flat fee, so that I, a lowly one-man crawler developer will not be able to afford it? Will it be that only Google can afford it, thus making their monopoly position even stronger?

Is there a Wikipedia crawling "welfare" program if I'm not a trillion dollar mega company?

Re: Only Google is really allowed to crawl the web

#323

Earlier quoted context omitted.

There is if you are doing it for work. For example, your company could get sued if you are found using that data and ignoring the ToS. If you are a public figure, you could get your name tarnished as doing something unethical or the media may call it "hacking". If you are rereleasing the data then you risk getting a takedown notice.

robots.txt is not a terms of service. Even if it was, it wouldn't be enforceable for a public website. You would need to prove that a web crawler is maliciously causing disruption to your service, and that is not easy.

All it takes is your company execs or lawyers to be afraid of a stern letter, and ask you to cancel your project. If you're violating their robots.txt, you're probably violating their terms of service that's hidden somewhere. And your company doesn't want to risk having to pay hundreds of thousands to fight a court case. There's also venues besides courts for them to attack you, like contacting the publishers or hosting platforms for your derivative works. It's a chilling effect.

And I'm not making this up. This kind of stuff has happened to me many times.

Re: Only Google is really allowed to crawl the web

#324
post #311

Earlier quoted context omitted.

On the other hand, I do not want my site to go down thanks to a few bad 'crawlers' that fork() a thousand http requests every second and take down my site, forcing me to do manual blocking or pay for a bigger server/scale-out my infrastructure. Why should I have to serve them?

Agree... this entire argument seems to think I give a rats ass about 9000 different crawlers that give me literally zero benefit and only waste server resources. Most of those crawlers are for ad-soaked piss poor search engines. I'd rather just block them all and allow crawlers that don't know how to behave.

[deleted]

Re: Only Google is really allowed to crawl the web

#325

Earlier quoted context omitted.

Even before it gets to that point, they routinely display snippets off regular websites and show ads next to it. Keeping users from clicking through to organic results helps them generate more revenue.

Here’a a thought: most companies don’t actually want to serve a ton of extra pages. For example, airlines just want to fly passengers. They don’t care who puts those butts in seats and they would fully acknowledge that they aren’t able to deliver a better flight search than Google can. I mean, sure, some small team of web developers at every airline is pissed, but the CEO needs butts in seats to keep the pilot and se…

"For example, airlines just want to fly passengers. They don’t care who puts those butts in seats"

Not sure who you talked to, but I've never heard that before. They all want to sell more direct and forego GDS fees and/or other types of fees and commissions. I'd love to see a quote from an airline VP or above that they don't care about their distribution model, boosting direct sales percentages, etc.

Re: Only Google is really allowed to crawl the web

#327

Earlier quoted context omitted.

Where's the money?

That's like arguing that newspapers are not a market, because it makes money from ads.

No, I was asking if web search was a market independent of ads.

BTW, newspapers also make money from subscriptions and sales of copies, so your analogy is doubly wrong.

Re: Only Google is really allowed to crawl the web

#329
post #237

Earlier quoted context omitted.

consumers are in this case the advertisers. google has a monopoly on search ads and does enforce it, being a drain on the economy since in many fields you only succeed if you spend on search ads

googles answer to this at yesterdays hearing.. Search isnt a single category. If you break it down, they arent a monopoly. For example. 1/2 of PRODUCT SEARCHES begin on Amazon. It's probably hard to argue Google as a monopoly if who they see as their main competitor has half the market share.

Anecdotally, its true for books too. Amazon is a great way to figure out which books are the best reviewed before deciding to get one (whether paper, kindle or 'other means').

Re: Only Google is really allowed to crawl the web

#330

Earlier quoted context omitted.

Even before it gets to that point, they routinely display snippets off regular websites and show ads next to it. Keeping users from clicking through to organic results helps them generate more revenue.

Here’a a thought: most companies don’t actually want to serve a ton of extra pages. For example, airlines just want to fly passengers. They don’t care who puts those butts in seats and they would fully acknowledge that they aren’t able to deliver a better flight search than Google can. I mean, sure, some small team of web developers at every airline is pissed, but the CEO needs butts in seats to keep the pilot and se…

the OP wasn't talking about google competing with airlines, but with other flight aggregating and search/booking services, by abusing their monopoly on web search.
Post reply on HN