Live data from Hacker News

Google Removed 749M Anna's Archive URLs from Its Search Results

torrentfreak.com

111–120 of 168 posts

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#112

Earlier quoted context omitted.

You can turn off personalization. (Operating under the assumption that most people search for facts, I personally don't see why one would ever want personalized results.)

Location based personalization is pretty useful - if I search for 'Bob's Discount Linguine' I want the one in my neighborhood. Lots of niche things (like programming) also reuse common english words to mean specific things - if I search e.g. 'locking' it's nice to get results related to asynchronous programming instead of locksmiths because google knows I regularly search for programming related terminology. Of cours…

For the better part of a decade it seems that every verb or noun I search for, all the top search results are some movie or TV show named after that verb or noun. And I've watched exactly two movies in the past two decades (Star Wars VII when it came out, and Alien just last week).

Sometimes I consider actually enabling personalized search just to get to the things that I'm actually looking for.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#113
post #96

Earlier quoted context omitted.

Publicly traded corporations are machines whose only lawful purpose is to make money. They are legally obligated to be sociopathic systems. They aren't evil like an axe murderer, they're evil like a gasoline fire. They may be useful when properly controlled, but they're certainly never worth defending in the way you seem to feel the need to

>Publicly traded corporations are machines whose only lawful purpose is to make money. Hey, so this isn't the case at all, publicly traded companies are under no lawful obligation to focus only on making money. Fiduciary duty does not mean this in any way. It's a common misconception whose perpetuation is harmful. Let's stop doing it.

Potato, potahto. While you're right that the law doesn't state it, it's also true that it is the only goal they have, so there's that.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#114

Earlier quoted context omitted.

I won't bother defending Google-style personalization as it exists for their search results, but since collisions in terminology across fields are common, it's not that hard to see how actual, thoughtful personalization could be useful. Someone searching for "Kafka" is going to want very different results based on whether they're thinking of software or literature. Opinions may also differ over the usefulness of sour…

> Kafka" is going to want very different results based on whether they're thinking of software or literature. Speak for yourself. I've worked in several "Kafka-esque" software organizations.

Arguably Google SERPs are getting closer to The Trial.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#115

Earlier quoted context omitted.

Google only ever returns a maximum of Bing only returns 900. Kagi only 200. Deep search and surfing is pretty much gone on all major search "engines".

> Google only ever returns a maximum of That's perfectly fine. If I'm going to use a search engine, I'm not willing to sift through hundreds of potentially relevant results. I hope I find what I'm searching for in the first page, or at best in the first 3 pages or so. What's not cool about Google is that now it hits you with AI slop with dubious quality right at the top, followed by a page of sponsored results, follo…

It's not Google's fault alone.

SEO manipulation for example, that could be tackled by our legal system similar to existing slander, unfair competition and advertising regulations. But unfortunately, most representatives are not digital natives but old digital buffoons, and the post-2000/Gen Z kids never gained an understanding of what actually makes the web tick.

As for the TLD explosion, we definitely need a completely new setup for ICANN. The trouble all of that has caused, just for a measly 250k in fees for each new gTLD, is insane.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#116

Earlier quoted context omitted.

Edit: after the 3rd page Source: https://www.courtlistener.com/docket/18552824/1436/united-st... For fun what Gemini says: “The notion that Google explicitly admitted to "deprioritizing good results to sell more ads" is a common interpretation of these documents and expert testimony.”

> Source: https://www.courtlistener.com/docket/18552824/1436/united-st... That's a 230-page pdf. Do you have a more specific citation?

I passed the PDF to Claude and asked it to check if there is any part of the document that states that google deprioritizes good search results in favor of advertisement. Here is the output from Claude:

Yes, the document contains highly significant factual findings by the Court regarding how Google deprioritized organic search results in favor of advertising. The most significant findings: The Court documents that the positioning of Google's AI features (AI Overviews, WebAnswers) on the search results page reduced users' interactions with organic web results - deliberately.

Relevant text:

"Some evidence suggests that placement of features like AI Overviews on the SERP has reduced user interactions with organic web results (i.e., the traditional "10 blue links")."

And:

"Placement of features like AI Overviews on the SERP has reduced user interactions with organic web results where Google's WebAnswers appears on the SERP"

Important note: these are not "admissions" in the sense of Google voluntarily confessing, but rather factual findings by the Court based on evidence presented during the trial - which is legally even more binding.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#118

Feels weird to say but I have found using Yandex of all places an excellent search engine for content that get taken down by DMCA requests. Eg if you want to watch a movie that's not on Netflix using a web stream the search results are far better. Feels like Google circa 2005.

This has been my search engine quality test for quite some time.

A good search engine will show you pirate websites because they have a comprehensive index. A great search engine will put them at the top of the list ahead of the fake results.

A great search engine that endures long enough attracts the type of attention that forces them to delist those results. Once you can no longer find that type of results you know it's time to look somewhere else.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#119
post #21

Feels weird to say but I have found using Yandex of all places an excellent search engine for content that get taken down by DMCA requests. Eg if you want to watch a movie that's not on Netflix using a web stream the search results are far better. Feels like Google circa 2005.

I've been playing around with a variety of search engines such as Kagi, Startpage, Ecosia, DDG. All of them are better than google in finding relevant results. Lol Google is way too "personalized".

The fact that Google seemingly returns results worse than Kagi, Startpage and Ecosia is just strange, given that Google provides search results for all three of them. Both Kagi and Ecosia uses other sources as well, I don't know about Startpage, so that's certainly part of it, but it still feels a little strange.

From using Ecosia, DuckDuckGo and Bing, I'd also argue that Bing is simply a better search engine at this point.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#120
post #83

Earlier quoted context omitted.

I was more suggesting that I want my LLM provider to launder the IP so it avoids copyright law. The LLM provider is a fancy search engine where copyright does not apply to the results.

Do LLMs filter piracy requests? For example, how will it respond to 'find me a free copy of the Lord of the Rings movies' or more explicitly 'find me a pirated copy ...'?

> how will it respond to 'find me a free copy of the Lord of the Rings movies' or more explicitly 'find me a pirated copy ...'

Apparently it depends on the model. Testing on OpenRouter with Search enabled, gpt-5 strictly refuses to provide any links, but Deepseek R1 provides several Archive.org links, one of which is for a torrent file.

Thanks Deepseek, I guess I'll be watching The Fellowship of The King for free tonight. ;)

Post reply on HN