Live data from Hacker News

Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

techdirt.com

101–110 of 164 posts

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#101
post #9

If only GPT wouldn't refuse my requests to write a crawler for $site. :(

Gemini has been perfectly willing to write such things for me

Grok Build won't even question it. You gotta weigh that Grok use up, though.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#102
Missing from this blog post is that the suit was filed because OpenAI was using SerpApi to collect Google search results

Alphabet is an Anthropic investor

Looking forward to the Amended Complaint by August 10

https://searchengineland.com/inside-google-searchguard-46767...

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#103
post #42

Earlier quoted context omitted.

I recall a articale I read "somewhere" that reported the Facebook makes big profit (Billions) from scams. So they have no ( or no strong motive) motive to shut such scams down An example is courtcase against Meta for using a Australians mining billionares likeness to promote a crypo investment scam. https://www.afr.com/technology/dad-it-s-a-fraud-call-that-sp...

I'm surprised legitimate companies don't pressure Facebook on this though. There are enough scams on Facebook that I now refuse to believe anything there, even though some of the things look useful and probably are not scams (and also are things I didn't know existed without an ad - thus filling one of the legitimate values of advertisements: informing me of things that would make my life better but I don't know exis…

Anyone making a Kickstarter knows that within days if not hours all your images, pictures, renderings will be harvested and dozens of "copycat" sites will be "selling" your product now, regardless of whether its a real thing or not yet, and they'll be advertising it on FB and IG.

I say those words in quotes - they have no intention of shipping you anything, just skimming low-hanging fruit from someone's ideas.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#104
post #89

Earlier quoted context omitted.

Almost certainly. That's OK, isn't it? Folks have been Googling things poorly for as long as there has been a Google to Google with, and now they have bots that do it on their behalf. The only issue is that the eyeballs stayed with the LLM, which allows it to hold a position that is potentially very powerful. This is particularly problematic with people who believe that computers are infallible.

The point of this thread is that you cannot replace search engines with an LLM. If your example is an LLM using a search engine as a tool call, you have not replaced the search engine. It’s still there.

Yes. An external search engine is still there, for now.

But the user doesn't necessarily know anything about that, and they don't necessarily care.

If/when the time is reached when external search engines are no longer present in the LLM loop, it seems likely that regular folks won't even notice this shift. As long as the answer-making machine keeps making answers, they won't have any reason to pay attention to this kind of back-end minutiae at all.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#105

(IANAL) I think that the deeper thing from this lawsuit is that from my understanding, (inherently) Search engines are considered public indexes and the data (URL's,index) behind it is considered uncopyrighted and as such aren't protected by DMCA because DMCA only works for copyrighted contents and thus the dismissal of the lawsuit by the Judge. Basically, search engines are publicly scrapable, though I do wonder as…

In that case, even if using a service violates its terms of service, distilling its AI output is not necessarily illegal under the law.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#106

Earlier quoted context omitted.

Google crawling respects robots.txt and doesn't break capcha. It is easy to tell Google to piss off. SerpAPI fully relies on end-user proxies distributed like malware (in LG tv apps for instance), it has no other way it could function because it exclusively ingests data from sources that tell it to stop. If you wanted to scrape a bot friendly site, you wouldn't need SerpApi

I'm vaguely sympathetic to this argument. But only vaguely. Google uses its monopoly position in advertising to basically ensure that you allow them to scrape your site (or if not you personally, the majority of revenue driving sites). They have the benefit of being allowed by default. They also then scrape again at the user level for users operating chrome. They also conveniently ignore global blocks for their adsbo…

And they do specifically also scrape sites anonymously, ostensibly to ensure you don't serve different content to GoogleBot than users (though this is in my experience unreliable at best, and that's even before we get to the "do we trust that that's the only scope?").

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#107
post #89

Earlier quoted context omitted.

Almost certainly. That's OK, isn't it? Folks have been Googling things poorly for as long as there has been a Google to Google with, and now they have bots that do it on their behalf. The only issue is that the eyeballs stayed with the LLM, which allows it to hold a position that is potentially very powerful. This is particularly problematic with people who believe that computers are infallible.

The point of this thread is that you cannot replace search engines with an LLM. If your example is an LLM using a search engine as a tool call, you have not replaced the search engine. It’s still there.

you have changed the monetization model for the search engine though

and maybe you really need the index and not the engine?

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#108

I'd be more OK with this if Google had a good API for their search results. But they've deprecated it, and now there is no alternative. So I'll continue to use 3rd parties that scrape Google results, until they change their mind.

I stopped using Google once it started requiring JS for search. That was the day the open web truly died.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#109
post #25

Earlier quoted context omitted.

It's interesting to imagine how the world would be different, if internet advertising giants were partially liable for scams/malware that they facilitate.

Ironically it's only really people who heavily ad-block/privacy-protect that get these scam/malware ads. When Google has no profile on you, your view is virtually worthless, so it's only bottom feeders that bid on those views. Average users get Coke and Tide ads. Its usually the most technically adept that get the worst ads, and usually they just turn their ad block back on.

From where did you come to conclusion that adblock users get malware ads? This is the first time I am hearing this thing.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#110

I'd be more OK with this if Google had a good API for their search results. But they've deprecated it, and now there is no alternative. So I'll continue to use 3rd parties that scrape Google results, until they change their mind.

I stopped using Google once it started requiring JS for search. That was the day the open web truly died.

EMCAScript has an open standard that anyone can implement.
Post reply on HN