Live data from Hacker News

Google Removed 749M Anna's Archive URLs from Its Search Results

torrentfreak.com

71–80 of 168 posts

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#72

Earlier quoted context omitted.

You can turn off personalization. (Operating under the assumption that most people search for facts, I personally don't see why one would ever want personalized results.)

Location based personalization is pretty useful - if I search for 'Bob's Discount Linguine' I want the one in my neighborhood. Lots of niche things (like programming) also reuse common english words to mean specific things - if I search e.g. 'locking' it's nice to get results related to asynchronous programming instead of locksmiths because google knows I regularly search for programming related terminology. Of cours…

Search results are still location-specific even if you disable personalization.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#73

Man I need to get around to downloading the z-archive torrents before annas archive is taken down. If I eliminate large PDFs and non english books I think I can fit it on two 32 TB drives with BTRFS z-std compression max setting. https://annas-archive.org/torrents

Let me know of those efforts, I wanna have an English/German/French backup of the archive, too. But as you said HDDs and filesystems are the problem, really.

Maybe I'll have to build a torrent splitter or something, because the UIs of all torrent clients are just not built for that.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#74
post #59

Earlier quoted context omitted.

They’re… yes. Yes, that’s exactly what they have done and continue to do. Are you familiar with it?

I think the comment is saying Google was also doing that.

They're doing one site less now

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#75
post #59

Earlier quoted context omitted.

They’re… yes. Yes, that’s exactly what they have done and continue to do. Are you familiar with it?

I think the comment is saying Google was also doing that.

Anna's archive doesn't engage in privacy-eroding antitrust/monopolistic activities (yet), so there's that I suppose...

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#76

Earlier quoted context omitted.

Google hides the most relevant results on the 3rd page. It was confirmed in trial disclosures a few months ago. Their concern isn’t public search.

Google only ever returns a maximum of Bing only returns 900. Kagi only 200. Deep search and surfing is pretty much gone on all major search "engines".

> Google only ever returns a maximum of That's perfectly fine. If I'm going to use a search engine, I'm not willing to sift through hundreds of potentially relevant results. I hope I find what I'm searching for in the first page, or at best in the first 3 pages or so.

What's not cool about Google is that now it hits you with AI slop with dubious quality right at the top, followed by a page of sponsored results, followed by some potentially useful results, followed by an entire ocean of spam traps and clone sites and really shady results with exotic never-seen-before TLDs that leaves you wondering whether clicking on a link will get in a hostile database. That's what's not cool about Google: is that you can't use it to search the web anymore.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#77

Earlier quoted context omitted.

1. Your chatbot doesn't have its own internet scale search index. 2. You're being given information that may or may not be coming in part from junk sites. All you've done is give up the agency to look at sources and decide for yourself which ones are legitimate.

I’m quite happy trading off the agency of wading through trash to an LLM. In fact, I would say that’s something they’re pretty good at.

It’s just regurgitating the same trash to you though.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#78
post #21

Feels weird to say but I have found using Yandex of all places an excellent search engine for content that get taken down by DMCA requests. Eg if you want to watch a movie that's not on Netflix using a web stream the search results are far better. Feels like Google circa 2005.

I've been playing around with a variety of search engines such as Kagi, Startpage, Ecosia, DDG. All of them are better than google in finding relevant results. Lol Google is way too "personalized".

I switched to Kagi a while back and ended up buying their annual subscription for unlimited searches. It's such a breath of fresh air, like a search engine from an alternate universe where Google just focused on search instead of adtech.

Re: Google Removed 749M Anna's Archive URLs from Its Search Results

#80

Man I need to get around to downloading the z-archive torrents before annas archive is taken down. If I eliminate large PDFs and non english books I think I can fit it on two 32 TB drives with BTRFS z-std compression max setting. https://annas-archive.org/torrents

> eliminate large PDFs

How large? Isn't that going to result in an arbitrary filter of books? In other domains, large PDFs are due to PDF production errors, such as using color or needlessly high resolution, and not so much due to the volume of content - at least for text.

Post reply on HN