Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

351–360 of 365 posts

Re: Only Google is really allowed to crawl the web

#351

Earlier quoted context omitted.

Aren't there anti trust laws to prevent this kind of thing?

The current anti-trust doctrine in the US has a goal of protecting consumers - not competition. What Google is doing is arguably great for consumers but awful to their competitors/other organizations. Technically, companies can simply block Google using robots.txt - but in reality that will lose them more money than the current partial disintermediation by Google is costing them - and Google knows this. It's a tall o…

Which is shortsighted. If competition did not benefit consumers, there would be no need for it anyway.

Re: Only Google is really allowed to crawl the web

#352

On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…

I wonder what happens to RSS feeds in this situation. Programs I run that process RSS feeds will just fetch them over HTTP completely headlessly, so if there are any CAPTCHAs, I'm not going to see them.

I've found that Cloudflare isn't great at this. I even found cases where my site was failing to load to googlebot (a "good" bot that they probably have the IPs for) because they were serving a captcha instead of my CSS.

So your best bet is setting a page rule to allow those URLs.

Re: Only Google is really allowed to crawl the web

#353

Earlier quoted context omitted.

> Isn't that the website owners right though? No. The internet is public. Publishers shouldn't get any say in who accesses their content or how they do it. As far as I'm concerned, the fact that they do is a bug.

No, it's not. I can setup a login page and keep you out if I want. And I can do it however I want.

But your login page will be public and subject to being crawled.

Re: Only Google is really allowed to crawl the web

#354

Earlier quoted context omitted.

Just tell people to stop using google. Go direct.

Upvoted - regardless how pointless some people might think this comment is, it really is the ONLY way that Google is going to drop out of its aggregate lead position. Enough people realizing Google is trapping and cannibalizing traffic to the other sites it feeds off of, and choosing to do other things EXCEPT touching Google properties, is THE ONLY way they'll be unseated. No clear legal path to stop a bully means it…

I find those little snippets actually mostly worthless, maybe because I’ve seen enough of them taken out of context or basically using a snippet from someone who figured out SEO properly, meanwhile the correct information may be down a couple links or not there at all.

Re: Only Google is really allowed to crawl the web

#355
post #237

Earlier quoted context omitted.

consumers are in this case the advertisers. google has a monopoly on search ads and does enforce it, being a drain on the economy since in many fields you only succeed if you spend on search ads

googles answer to this at yesterdays hearing.. Search isnt a single category. If you break it down, they arent a monopoly. For example. 1/2 of PRODUCT SEARCHES begin on Amazon. It's probably hard to argue Google as a monopoly if who they see as their main competitor has half the market share.

The US is so behind in identifying markets in technology which is what is leading to this dominance by a few companies and their resulting monopoly like power. We had already figured out that you can be dominant in only a subset of a market. For example Disney was forced to sell off fox sports channels when it purchased fox because it already owned espn and would have dominated sports TV. That’s the thing it wasn’t even just TV but a subset, sports TV. That identification is where we are behind. As of now no, one in the FTC knows what makes a market or why say YouTube and Facebook may both show large amounts of video content but are absolutely not competitors in video content space. It is because the functions are completely different. YouTube is barely a social network at all despite having users and comments and pages. Facebook is hardly a video platform at all because it isn’t profitable for users to focus on Facebook videos and make ad money.

Re: Only Google is really allowed to crawl the web

#356

Earlier quoted context omitted.

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

Aren't there anti trust laws to prevent this kind of thing?

Laws are one thing and enforceability another.

In fact Google had to make certain concessions in order for the Google Flights acquisition to get regulatory approval.

IIRC a Chinese firewall between Google data and Google Flights...but like many regulations they were likely written by Google lobbyists aka the industry experts. Because at the end of the day Google flights: 1. Still has the built in widget above organic results and 2. They still bid on their own ad spots jacking up costs on competition which is ultimately passed on to consumers.

Re: Only Google is really allowed to crawl the web

#357
post #301

Earlier quoted context omitted.

Fine, the content is free. But if your crawlers want access to the content, then pay! Simple as that.

Will it be a flat fee, so that I, a lowly one-man crawler developer will not be able to afford it? Will it be that only Google can afford it, thus making their monopoly position even stronger? Is there a Wikipedia crawling "welfare" program if I'm not a trillion dollar mega company?

Sure! Apply to become a crawler. And if you meet certain criteria and your crawlers don’t exceed a quota then have at it. The key is not to make it technically challenging, but to erect a legal barrier.

Re: Only Google is really allowed to crawl the web

#359

Earlier quoted context omitted.

Aren't there anti trust laws to prevent this kind of thing?

The current anti-trust doctrine in the US has a goal of protecting consumers - not competition. What Google is doing is arguably great for consumers but awful to their competitors/other organizations. Technically, companies can simply block Google using robots.txt - but in reality that will lose them more money than the current partial disintermediation by Google is costing them - and Google knows this. It's a tall o…

Anti-trust above all recognizes competition benefits consumers. And so “unfairly competing” is prohibited, because it is bad for the market, thus bad for consumers.

Re: Only Google is really allowed to crawl the web

#360
post #345
post #254

Earlier quoted context omitted.

I can say this will give a lot of businesses false sense of security. It is already bypassable. the Web scraping technology that I am aware of has reached end game already: Unless you are prepared to authenticate every user/visitor to your website with a dollar sign, lobby congress to pass a bill to outlaw web scraping, you will not be able to stop web scraping in 2021 and beyond.

But what about captchas? Due to aggressive no-script and uBlock use I, browsing the website as a human, keep getting hit by captchas and my success rate is falling to a coinflip. If there's a script to automate that I'm all ears.

100% doable. Like I said these type of blanket throttling seems to be the new trend but it's already defeated.

I just no longer see it possible to 1) put information on the web (private or public) 2) give access outside your organization (customers or visitors) 3) expect your website will not be scraped.

ToS is NOT the law unfortunately.

Post reply on HN