Live data from Hacker News

Why we're taking legal action against SerpApi's unlawful scraping

blog.google

51–60 of 113 posts

Re: Why we're taking legal action against SerpApi's unlawful scraping

#51
post #41

Earlier quoted context omitted.

Because they've forced everyone to allow them. They're the internet traffic mafia. Block them and you disappear from the internet They abuse this power to scrape your work, summarize it and cut you out as much as possible. Pure value extraction of others' work without equal return. Now intensified with AI But yeah, you're right. They're not deceptive

> Because they've forced everyone to allow them. nobody is forcing anyone. This is the same argument that people said about google search. Nobody is forcing anyone to use google search, google chrome, or even allow googlebot for scraping. Thousands of poeple have switched over to chatgpt, brave/firefox .. Your argument sounds like "I dont like Apple's practices, and I'm forced to buy iPhones. No buddy, if you dont li…

> Thousands of poeple have switched over to chatgpt, brave/firefox ..

If you want people to visit your website, limiting yourself to the "thousands" of people who don't use google isn't really an option.

> Your argument sounds like "I dont like Apple's practices, and I'm forced to buy iPhones. No buddy, if you dont like Apple, dont buy their products"

Well, I don't like Apple's or Google's practices, but I basically [1] have to use either iOS or Android.

[1]: yes there are things like GrapheneOS and librem, but those aren't really practical for most people.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#52
post #7
post #3

I'm not sure of the legality but I definitely appreciate their product. This lawsuit seems odd because google themselves scrape content for their indexes. From what I see SerpApi is really just providing a machine interface that Google themselves refuses to provide users and visibility into SERPs which is also something that users should have available to them. I'm probably just being naive though...

Google publishes how to control their bot - with robots.txt. They then obey those instructions. Google also takes some effort to not use all your bandwidth. Google isn't perfect, but they are at least making a "good faith" effort to be nice and this does count in court. Overall most will agree that in general what google does to allow people to find their website is worth the things that google is doing. You can of c…

Who says robots.txt is legally binding? Where's the Sherman Antitrust analysis?I'm more confused than before.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#53
post #15

Earlier quoted context omitted.

There's no law that says you have to do that. It used to be a sensible thing to do, in the early internet. In the current internet, obeying robots.txt is a self-handicap and you shouldn't do it. DDoS remains illegal regardless of robots.txt.

It's rather odd to use words like "should" when you're advocating for disrespecting other people's wishes. There are sometimes reasons not to cooperate, but it seems like a good default.

The web is now hostile. If you're starting a search engine, everyone else has written a robots.txt that bans you from starting a search engine. You either ignore that, or you abandon your plan to make a search engine.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#54
post #7

Earlier quoted context omitted.

Google publishes how to control their bot - with robots.txt. They then obey those instructions. Google also takes some effort to not use all your bandwidth. Google isn't perfect, but they are at least making a "good faith" effort to be nice and this does count in court. Overall most will agree that in general what google does to allow people to find their website is worth the things that google is doing. You can of c…

Who says robots.txt is legally binding? Where's the Sherman Antitrust analysis?I'm more confused than before.

The courts say. With this as a long standing tradition they are likely to agree.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#55
post #53

Earlier quoted context omitted.

It's rather odd to use words like "should" when you're advocating for disrespecting other people's wishes. There are sometimes reasons not to cooperate, but it seems like a good default.

The web is now hostile. If you're starting a search engine, everyone else has written a robots.txt that bans you from starting a search engine. You either ignore that, or you abandon your plan to make a search engine.

Maybe only ethical choice is not to play? Or to do it the hard way. Scrape what people allow and try to make deals to get more data.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#56

SerpApi wouldn't even be a thing if Google offered an equivalent API...

why does google need to offer it?

Because Google scrapes other site's data to build its AI market dominance in Gemini. The promise of web 2.0 was APIs, Google aims to cement its position in web 4.0 while suing others for doing what it does on a mass scale.

Adversarial Interoperability is Digital Human Right. Either companies can provide it reasonably or the people will assert their rights through other means.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#57
post #53

Earlier quoted context omitted.

The web is now hostile. If you're starting a search engine, everyone else has written a robots.txt that bans you from starting a search engine. You either ignore that, or you abandon your plan to make a search engine.

Maybe only ethical choice is not to play? Or to do it the hard way. Scrape what people allow and try to make deals to get more data.

"Making a search engine is unethical" is certainly one of the takes of all time. I'm sure Google is glad you believe it.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#58
post #39
post #28

Earlier quoted context omitted.

What about for their LLM products? We know that OpenAi does not respect the robots.txt file

Google uses the same crawler and robots.txt file for training data.

It's actually a different crawler for training data: Googlebot-extended so you can exclude yourself from the training data though not the search summaries.

Re: Why we're taking legal action against SerpApi's unlawful scraping

#60
post #42

Earlier quoted context omitted.

At this point everyone knows about robots.txt, so if you didn't opt-out that is your own fault. Opting out of everyone at once is easy, and you get fine grained control if you want it. Also most people would agree they are fine with being indexed in general. That is different from email spam where people don't want it.

Looking at SerpApi clients, looks like most companies would agree they are fine with scraping Google. That is different from having your website content stolen and summarized by AI on Google search, which people don't want.

Or by Google codewiki, which is morally the equivalent to making a business out of ersatz travel guides by ripping off the authors of real ones
Post reply on HN