Live data from Hacker News

Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

techdirt.com

91–100 of 164 posts

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#91
post #77

Earlier quoted context omitted.

> is a problem when it is offered without context, sourcing, etc. Sure. But that's not what's happening here, yeah? Both examples are providing sources. Actually I'd say that the LLM is doing a better job at that. You click a button and see 9 URLs to 9 news sites. I think for the "influence LLM results" fears, it would be ridiculous to argue about which multi-billion dollar American company can be trusted more not to…

> Actually I'd say that the LLM is doing a better job at that. You click a button and see 9 URLs to 9 news sites. Does it show relevant snippets for each of the 9 URLs? If not, then it's as bad as Google's AI summary. It's not that rare that the reference disagrees with the LLM. The value of the non-LLM search engine is that in the search results, you see the relevant snippet, and if you want more information, you kn…

> Does it show relevant snippets for each of the 9 URLs? If not, then it's as bad as Google's AI summary.

Sounds like you're saying that the LLM isn't any better than Google's LLM-created snippets. The claim was that Google has no moat against LLMs, and your evidence against this claim is Google's LLM generated content.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#92
post #69

Earlier quoted context omitted.

Imagine how world would be different if the actual scammers and malware creators would be prosecuted and not insteead demanded that internet giants play a privatized police force.

Aren't the giants closer to the mob than a police force? I suppose those things are not terribly different in practice, but Google, Facebook, et al make money from leaving the scams on their ad networks.

Either way works - I'd still prefer the scammers to be liable and persecuted for their scamming than demanding that platforms enact censorship and policeing control.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#93

Earlier quoted context omitted.

Let me tell you about this new little-known technology that's been gaining traction in the last few years...

Sure, go ahead and ask the LLM who won the 2026 World Cup. Or nearby BBQ places open now. Even assuming no hallucinations, an LLM can replace some use cases for search but not all.

On a personal level if I used to do 100 google searches on a daily basis, now I would be doing only 10. And it is the case with everyone in the tech ecosystem at least. So it would be fair to assume that there share has reduced.

On the LLM side as well, I have not seen much people using Gemini vs the market share of Claude / Open AI.

The assumption of the long tail still using Gemini because it is bundled might be correct, but I am not even sure if that is something that Google will be happy with.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#94

EU protects a database creator if there has been a qualitative or quantitative "substantial investment" in obtaining, verifying, or presenting the content, regardless of creative expression. In USA copyright requires a minimum degree of original creativity in the selection, coordination, or arrangement of the data. I think it's a rather grey line to say that Google search results are just facts, but eg maps are copyr…

> I think it's a rather grey line to say that Google search results are just facts, but eg maps are copyrightable. I don't think a map is a good example. A picture is probably a better one. A map, by its very definition, is not a replica of any part of the original artifact. Not merely because a map is not the territory, but also because it's not even a direct, unaltered view of the thing. There is clearly some creat…

[deleted]

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#96
post #74

Earlier quoted context omitted.

The model just asked the crawler to get a recent copy of relevant pages and used that to give the answer. That is why AI companies experimenting with browsers or at least agentic extensions to browsers as that allows to fetch on behalf of the user directly from the device with a residential IP and ditch expensive to maintain crawling infrastructure.

> relevant pages If you want to know what the relevant pages are, you need a search index.

Index is just a part of the model. The nice property is that the index for LLM does not need to be updated often so LLM plus index can be static. This even works for the latest news as LLM can fetch major news sites to learn the latest stories and then fetch individual articles for the final answer.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#97
post #77

Earlier quoted context omitted.

> Actually I'd say that the LLM is doing a better job at that. You click a button and see 9 URLs to 9 news sites. Does it show relevant snippets for each of the 9 URLs? If not, then it's as bad as Google's AI summary. It's not that rare that the reference disagrees with the LLM. The value of the non-LLM search engine is that in the search results, you see the relevant snippet, and if you want more information, you kn…

> Does it show relevant snippets for each of the 9 URLs? If not, then it's as bad as Google's AI summary. Sounds like you're saying that the LLM isn't any better than Google's LLM-created snippets. The claim was that Google has no moat against LLMs, and your evidence against this claim is Google's LLM generated content.

The discussion is about Google Search results vs LLM services (summary or otherwise).

So what I'm saying is:

Google Search Results > LLM results

And yes, that includes:

Google Search Results > Google Summary on search results page

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#98
post #16

This ruling might feel good viscerally, but it also reinforces Googles own scraping as perfectly legal. At its inception, Google probably viewed this lawsuit as win-win. Either they successfully sue a competitor into oblivion or establish a precedent that will protect themselves in the future. Google lost, but they still won.

> but it also reinforces Googles own scraping as perfectly legal.

I don't like Google very much, but making it illegal to scrape public data enables way too much abuse, so this is for the best.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#99
post #96

Earlier quoted context omitted.

> relevant pages If you want to know what the relevant pages are, you need a search index.

Index is just a part of the model. The nice property is that the index for LLM does not need to be updated often so LLM plus index can be static. This even works for the latest news as LLM can fetch major news sites to learn the latest stories and then fetch individual articles for the final answer.

> Index is just a part of the model

I can assure you that it is not. If you download any model off of huggingface, it does not also include an index of the internet.

That screenshot shows the model making a tool call to an external search index.

Re: Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped

#100
post #89
post #76

Earlier quoted context omitted.

Did it do a Google search to learn all that up to date information?

Almost certainly. That's OK, isn't it? Folks have been Googling things poorly for as long as there has been a Google to Google with, and now they have bots that do it on their behalf. The only issue is that the eyeballs stayed with the LLM, which allows it to hold a position that is potentially very powerful. This is particularly problematic with people who believe that computers are infallible.

The point of this thread is that you cannot replace search engines with an LLM. If your example is an LLM using a search engine as a tool call, you have not replaced the search engine. It’s still there.
Post reply on HN