What’s the difference between scraping and malicious scraping? Does google engage in scraping or malicious scraping? Do the AI companies engage in scraping or malicious scraping?
Why we're taking legal action against SerpApi's unlawful scraping
11–20 of 113 posts
Re: Why we're taking legal action against SerpApi's unlawful scraping
#12What’s the difference between scraping and malicious scraping? Does google engage in scraping or malicious scraping? Do the AI companies engage in scraping or malicious scraping?
Re: Why we're taking legal action against SerpApi's unlawful scraping
#13SerpApi wouldn't even be a thing if Google offered an equivalent API...
Re: Why we're taking legal action against SerpApi's unlawful scraping
#14And then pretending that they're fighting for other people's copyright is just the cherry on top of the pile of hypocrisy.
Re: Why we're taking legal action against SerpApi's unlawful scraping
#15> SerpApi deceptively takes content that Google licenses from others They have a different definition of "licensing" than most people I guess. Aren't site operators complaining about Google using this "licensed" content in AI overviews... not to mention the scraping for AI model training. The pot is calling the kettle black.
As far as I know, Google respects robots.txt and doesn't obfuscate their crawlers, so you can easily block them if you want. It seems like an important distinction?
DDoS remains illegal regardless of robots.txt.
Re: Why we're taking legal action against SerpApi's unlawful scraping
#16What’s the difference between scraping and malicious scraping? Does google engage in scraping or malicious scraping? Do the AI companies engage in scraping or malicious scraping?
> Stealthy scrapers like SerpApi override those directives and give sites no choice at all. SerpApi uses shady back doors — like cloaking themselves, bombarding websites with massive networks of bots and giving their crawlers fake and constantly changing names — circumventing our security measures to take websites’ content wholesale. [...] SerpApi deceptively takes content that Google licenses from others (like images that appear in Knowledge Panels, real-time data in Search features and much more), and then resells it for a fee. In doing so, it willfully disregards the rights and directives of websites and providers whose content appears in Search.
To me this seems... interesting, for sure. I think that Google already set a bad precedent by pulling content from the web directly into its results, and an even worse one by paying websites with user-generated content for said content (while those sites didn't pay the users that actually made the user-generated content, as an additional bitchslap.)
But it seems like at the very least Google is suggesting that SerpApi is effectively trying to "steal" the work Google did, rather than do the same work themselves. Though I wonder if this is really Google pulling up the ladder behind them a bit, given how privileged of a position they are in with regards to web scraping.
It's a tough case. I think that something does need to ultimately be done about "malicious" web scraping that ignores robots.txt, but traditionally that sort of thing did not violate any laws, and I feel somewhat skeptical that it will be found to violate the law today. I mean, didn't LinkedIn try this same thing?
Re: Why we're taking legal action against SerpApi's unlawful scraping
#17SerpApi wouldn't even be a thing if Google offered an equivalent API...
Re: Why we're taking legal action against SerpApi's unlawful scraping
#18What’s the difference between scraping and malicious scraping? Does google engage in scraping or malicious scraping? Do the AI companies engage in scraping or malicious scraping?
Re: Why we're taking legal action against SerpApi's unlawful scraping
#19Re: Why we're taking legal action against SerpApi's unlawful scraping
#20> SerpApi’s answer to SearchGuard is to mask the hundreds of millions of automated queries it is sending to Google each day to make them appear as if they are coming from human users. SerpApi’s founder recently described the process as “creating fake browsers using a multitude of IP addresses that Google sees as normal users.”