Live data from Hacker News

Why we're taking legal action against SerpApi's unlawful scraping

blog.google

41–50 of 113 posts

Re: Why we're taking legal action against SerpApi's unlawful scraping

#41
post #21

Earlier quoted context omitted.

No, Google doesn't use deceptive means. They identify their crawler as GoogleBot, and obey robots.txt.

Because they've forced everyone to allow them. They're the internet traffic mafia. Block them and you disappear from the internet They abuse this power to scrape your work, summarize it and cut you out as much as possible. Pure value extraction of others' work without equal return. Now intensified with AI But yeah, you're right. They're not deceptive

> Because they've forced everyone to allow them.

nobody is forcing anyone. This is the same argument that people said about google search. Nobody is forcing anyone to use google search, google chrome, or even allow googlebot for scraping.

Thousands of poeple have switched over to chatgpt, brave/firefox ..

Your argument sounds like "I dont like Apple's practices, and I'm forced to buy iPhones. No buddy, if you dont like Apple, dont buy their products"

Re: Why we're taking legal action against SerpApi's unlawful scraping

#42
post #7

Earlier quoted context omitted.

Google publishes how to control their bot - with robots.txt. They then obey those instructions. Google also takes some effort to not use all your bandwidth. Google isn't perfect, but they are at least making a "good faith" effort to be nice and this does count in court. Overall most will agree that in general what google does to allow people to find their website is worth the things that google is doing. You can of c…

But their robots are enabled by default. So it is a form of unsolicited scraping. If I spam millions of email addresses without asking for permission but provide a link to opt-out form, am I the good guy?

At this point everyone knows about robots.txt, so if you didn't opt-out that is your own fault. Opting out of everyone at once is easy, and you get fine grained control if you want it.

Also most people would agree they are fine with being indexed in general. That is different from email spam where people don't want it.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#43
post #41

Earlier quoted context omitted.

Because they've forced everyone to allow them. They're the internet traffic mafia. Block them and you disappear from the internet They abuse this power to scrape your work, summarize it and cut you out as much as possible. Pure value extraction of others' work without equal return. Now intensified with AI But yeah, you're right. They're not deceptive

> Because they've forced everyone to allow them. nobody is forcing anyone. This is the same argument that people said about google search. Nobody is forcing anyone to use google search, google chrome, or even allow googlebot for scraping. Thousands of poeple have switched over to chatgpt, brave/firefox .. Your argument sounds like "I dont like Apple's practices, and I'm forced to buy iPhones. No buddy, if you dont li…

> Your argument sounds like "I dont like Apple's practices, and I'm forced to buy iPhones. No buddy, if you dont like Apple, dont buy their products"

No, not really. There are alternatives to Apple. Whereas here Google controls the gate to the majority of internet traffic

For many it's "block Google and your business dies"

Re: Why we're taking legal action against SerpApi's unlawful scraping

#44
post #41

Earlier quoted context omitted.

Because they've forced everyone to allow them. They're the internet traffic mafia. Block them and you disappear from the internet They abuse this power to scrape your work, summarize it and cut you out as much as possible. Pure value extraction of others' work without equal return. Now intensified with AI But yeah, you're right. They're not deceptive

> Because they've forced everyone to allow them. nobody is forcing anyone. This is the same argument that people said about google search. Nobody is forcing anyone to use google search, google chrome, or even allow googlebot for scraping. Thousands of poeple have switched over to chatgpt, brave/firefox .. Your argument sounds like "I dont like Apple's practices, and I'm forced to buy iPhones. No buddy, if you dont li…

[deleted]

Re: Why we're taking legal action against SerpApi's unlawful scraping

#45
post #42

Earlier quoted context omitted.

But their robots are enabled by default. So it is a form of unsolicited scraping. If I spam millions of email addresses without asking for permission but provide a link to opt-out form, am I the good guy?

At this point everyone knows about robots.txt, so if you didn't opt-out that is your own fault. Opting out of everyone at once is easy, and you get fine grained control if you want it. Also most people would agree they are fine with being indexed in general. That is different from email spam where people don't want it.

Looking at SerpApi clients, looks like most companies would agree they are fine with scraping Google. That is different from having your website content stolen and summarized by AI on Google search, which people don't want.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#46
post #42

Earlier quoted context omitted.

At this point everyone knows about robots.txt, so if you didn't opt-out that is your own fault. Opting out of everyone at once is easy, and you get fine grained control if you want it. Also most people would agree they are fine with being indexed in general. That is different from email spam where people don't want it.

Looking at SerpApi clients, looks like most companies would agree they are fine with scraping Google. That is different from having your website content stolen and summarized by AI on Google search, which people don't want.

The claim is SerApi is not honoring robots.txt, and they are getting far more data from google/more often than needed for an index operation. Or at least that is the best I can make out of the claim in court from the article - I have not read the actual complaint.

People are generally fine with indexing operations so long as you don't use too much bandwidth.

Using AI to summarize content is still and open question - I wouldn't be surprised if this develops to some form of "you can index but not summarize", but only time will tell.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#48
post #46

Earlier quoted context omitted.

Looking at SerpApi clients, looks like most companies would agree they are fine with scraping Google. That is different from having your website content stolen and summarized by AI on Google search, which people don't want.

The claim is SerApi is not honoring robots.txt, and they are getting far more data from google/more often than needed for an index operation. Or at least that is the best I can make out of the claim in court from the article - I have not read the actual complaint. People are generally fine with indexing operations so long as you don't use too much bandwidth. Using AI to summarize content is still and open question -…

[deleted]

Re: Why we're taking legal action against SerpApi's unlawful scraping

#49
post #15
post #8

Earlier quoted context omitted.

As far as I know, Google respects robots.txt and doesn't obfuscate their crawlers, so you can easily block them if you want. It seems like an important distinction?

There's no law that says you have to do that. It used to be a sensible thing to do, in the early internet. In the current internet, obeying robots.txt is a self-handicap and you shouldn't do it. DDoS remains illegal regardless of robots.txt.

It's rather odd to use words like "should" when you're advocating for disrespecting other people's wishes. There are sometimes reasons not to cooperate, but it seems like a good default.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#50
post #8

Earlier quoted context omitted.

As far as I know, Google respects robots.txt and doesn't obfuscate their crawlers, so you can easily block them if you want. It seems like an important distinction?

Google can afford to respect robots.txt because it has a monopoly on search and nobody would consider actually blocking them in said robots.txt anyway. SerpApi doesn't have that privilege.

Some domains do block Google, often partially. There are some statistics here:

https://radar.cloudflare.com/ai-insights#ai-user-agents-foun...

Post reply on HN