Live data from Hacker News

Why we're taking legal action against SerpApi's unlawful scraping

blog.google

31–40 of 113 posts

Re: Why we're taking legal action against SerpApi's unlawful scraping

#31
post #21

> Defendant SerpApi, LLC (“SerpApi”) offers services that “scrape” this copyrighted content and more from Google, using deceptive means to automatically access and take it for free at an astonishing scale and then offering it to various customers for a fee. In doing so, SerpApi acquires for itself the valuable product of Google’s labors and investment in the content, and denies Google’s partners compensation for thei…

No, Google doesn't use deceptive means. They identify their crawler as GoogleBot, and obey robots.txt.

Because they've forced everyone to allow them. They're the internet traffic mafia. Block them and you disappear from the internet

They abuse this power to scrape your work, summarize it and cut you out as much as possible. Pure value extraction of others' work without equal return. Now intensified with AI

But yeah, you're right. They're not deceptive

Re: Why we're taking legal action against SerpApi's unlawful scraping

#32
post #7
post #3

I'm not sure of the legality but I definitely appreciate their product. This lawsuit seems odd because google themselves scrape content for their indexes. From what I see SerpApi is really just providing a machine interface that Google themselves refuses to provide users and visibility into SERPs which is also something that users should have available to them. I'm probably just being naive though...

Google publishes how to control their bot - with robots.txt. They then obey those instructions. Google also takes some effort to not use all your bandwidth. Google isn't perfect, but they are at least making a "good faith" effort to be nice and this does count in court. Overall most will agree that in general what google does to allow people to find their website is worth the things that google is doing. You can of c…

[deleted]

Re: Why we're taking legal action against SerpApi's unlawful scraping

#33
post #7
post #3

I'm not sure of the legality but I definitely appreciate their product. This lawsuit seems odd because google themselves scrape content for their indexes. From what I see SerpApi is really just providing a machine interface that Google themselves refuses to provide users and visibility into SERPs which is also something that users should have available to them. I'm probably just being naive though...

Google publishes how to control their bot - with robots.txt. They then obey those instructions. Google also takes some effort to not use all your bandwidth. Google isn't perfect, but they are at least making a "good faith" effort to be nice and this does count in court. Overall most will agree that in general what google does to allow people to find their website is worth the things that google is doing. You can of c…

But their robots are enabled by default. So it is a form of unsolicited scraping. If I spam millions of email addresses without asking for permission but provide a link to opt-out form, am I the good guy?

Re: Why we're taking legal action against SerpApi's unlawful scraping

#34
post #8
post #2

> SerpApi deceptively takes content that Google licenses from others They have a different definition of "licensing" than most people I guess. Aren't site operators complaining about Google using this "licensed" content in AI overviews... not to mention the scraping for AI model training. The pot is calling the kettle black.

As far as I know, Google respects robots.txt and doesn't obfuscate their crawlers, so you can easily block them if you want. It seems like an important distinction?

robots.txt is not a legally binding document, nobody needs to actually respect it

Re: Why we're taking legal action against SerpApi's unlawful scraping

#35
post #21

> Defendant SerpApi, LLC (“SerpApi”) offers services that “scrape” this copyrighted content and more from Google, using deceptive means to automatically access and take it for free at an astonishing scale and then offering it to various customers for a fee. In doing so, SerpApi acquires for itself the valuable product of Google’s labors and investment in the content, and denies Google’s partners compensation for thei…

No, Google doesn't use deceptive means. They identify their crawler as GoogleBot, and obey robots.txt.

[deleted]

Re: Why we're taking legal action against SerpApi's unlawful scraping

#36
post #8

Earlier quoted context omitted.

As far as I know, Google respects robots.txt and doesn't obfuscate their crawlers, so you can easily block them if you want. It seems like an important distinction?

Google can afford to respect robots.txt because it has a monopoly on search and nobody would consider actually blocking them in said robots.txt anyway. SerpApi doesn't have that privilege.

but SerpApi is not scraping websites, it is sending malicoius requests to google.com.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#37
post #36

Earlier quoted context omitted.

Google can afford to respect robots.txt because it has a monopoly on search and nobody would consider actually blocking them in said robots.txt anyway. SerpApi doesn't have that privilege.

but SerpApi is not scraping websites, it is sending malicoius requests to google.com.

SerpApi is scraping Google. The "maliciousness" if the requests is a matter of perspective. Of course Google considers it malicious; that doesn't necessarily make it true.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#38
Related:

Reddit Accuses 'Data Scraper' Companies of Stealing Its Information

https://news.ycombinator.com/item?id=45695433

Our Response to Reddit, Inc. vs. SerpApi, LLC: Defending the First Amendment

https://news.ycombinator.com/item?id=45739889

Re: Why we're taking legal action against SerpApi's unlawful scraping

#39
post #28
post #21

Earlier quoted context omitted.

No, Google doesn't use deceptive means. They identify their crawler as GoogleBot, and obey robots.txt.

What about for their LLM products? We know that OpenAi does not respect the robots.txt file

Google uses the same crawler and robots.txt file for training data.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#40
post #8

Earlier quoted context omitted.

As far as I know, Google respects robots.txt and doesn't obfuscate their crawlers, so you can easily block them if you want. It seems like an important distinction?

Google can afford to respect robots.txt because it has a monopoly on search and nobody would consider actually blocking them in said robots.txt anyway. SerpApi doesn't have that privilege.

Google has respected robots.txt from the start.
Post reply on HN