Live data from Hacker News

Why we're taking legal action against SerpApi's unlawful scraping

blog.google

11–20 of 113 posts

Re: Why we're taking legal action against SerpApi's unlawful scraping

#12

What’s the difference between scraping and malicious scraping? Does google engage in scraping or malicious scraping? Do the AI companies engage in scraping or malicious scraping?

Malicious scraping is when people other than them do it. When they scrape the internet to train their AI, it's "lawful" because they said so.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#15
post #8
post #2

> SerpApi deceptively takes content that Google licenses from others They have a different definition of "licensing" than most people I guess. Aren't site operators complaining about Google using this "licensed" content in AI overviews... not to mention the scraping for AI model training. The pot is calling the kettle black.

As far as I know, Google respects robots.txt and doesn't obfuscate their crawlers, so you can easily block them if you want. It seems like an important distinction?

There's no law that says you have to do that. It used to be a sensible thing to do, in the early internet. In the current internet, obeying robots.txt is a self-handicap and you shouldn't do it.

DDoS remains illegal regardless of robots.txt.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#16

What’s the difference between scraping and malicious scraping? Does google engage in scraping or malicious scraping? Do the AI companies engage in scraping or malicious scraping?

Note that I am not defending the merits of Google's lawsuit, but they did describe in this very post what they believe distinguishes their scraping versus SerpApi.

> Stealthy scrapers like SerpApi override those directives and give sites no choice at all. SerpApi uses shady back doors — like cloaking themselves, bombarding websites with massive networks of bots and giving their crawlers fake and constantly changing names — circumventing our security measures to take websites’ content wholesale. [...] SerpApi deceptively takes content that Google licenses from others (like images that appear in Knowledge Panels, real-time data in Search features and much more), and then resells it for a fee. In doing so, it willfully disregards the rights and directives of websites and providers whose content appears in Search.

To me this seems... interesting, for sure. I think that Google already set a bad precedent by pulling content from the web directly into its results, and an even worse one by paying websites with user-generated content for said content (while those sites didn't pay the users that actually made the user-generated content, as an additional bitchslap.)

But it seems like at the very least Google is suggesting that SerpApi is effectively trying to "steal" the work Google did, rather than do the same work themselves. Though I wonder if this is really Google pulling up the ladder behind them a bit, given how privileged of a position they are in with regards to web scraping.

It's a tough case. I think that something does need to ultimately be done about "malicious" web scraping that ignores robots.txt, but traditionally that sort of thing did not violate any laws, and I feel somewhat skeptical that it will be found to violate the law today. I mean, didn't LinkedIn try this same thing?

Re: Why we're taking legal action against SerpApi's unlawful scraping

#17

SerpApi wouldn't even be a thing if Google offered an equivalent API...

Why would Google offer an API? This is similar to saying when Apple sues an employee stealing IP "Nobody would steal the IP if they gave it away for free". The question is - why?

Re: Why we're taking legal action against SerpApi's unlawful scraping

#18

What’s the difference between scraping and malicious scraping? Does google engage in scraping or malicious scraping? Do the AI companies engage in scraping or malicious scraping?

Whether you obey robots.txt (Google does, SerpApi doesn't) seems like an important distinction.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#19
Google scrapes so what even is this? Beyond that I think it is unreasonable and monopolistic that Google can use all this data (like YouTube) to bolster their AI products but no one else can. It just means the megacorp will keep being megacorp and smaller players are doomed to have to work much harder and get very lucky. It’s not fair competition. So I view scraping Google as necessary for our society.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#20
From the filing

> SerpApi’s answer to SearchGuard is to mask the hundreds of millions of automated queries it is sending to Google each day to make them appear as if they are coming from human users. SerpApi’s founder recently described the process as “creating fake browsers using a multitude of IP addresses that Google sees as normal users.”

Post reply on HN