Google Removed 749M Anna's Archive URLs from Its Search Results
71–80 of 168 posts
Re: Google Removed 749M Anna's Archive URLs from Its Search Results
#72Earlier quoted context omitted.
You can turn off personalization. (Operating under the assumption that most people search for facts, I personally don't see why one would ever want personalized results.)
Location based personalization is pretty useful - if I search for 'Bob's Discount Linguine' I want the one in my neighborhood. Lots of niche things (like programming) also reuse common english words to mean specific things - if I search e.g. 'locking' it's nice to get results related to asynchronous programming instead of locksmiths because google knows I regularly search for programming related terminology. Of cours…
Re: Google Removed 749M Anna's Archive URLs from Its Search Results
#73Man I need to get around to downloading the z-archive torrents before annas archive is taken down. If I eliminate large PDFs and non english books I think I can fit it on two 32 TB drives with BTRFS z-std compression max setting. https://annas-archive.org/torrents
Maybe I'll have to build a torrent splitter or something, because the UIs of all torrent clients are just not built for that.
Re: Google Removed 749M Anna's Archive URLs from Its Search Results
#74Re: Google Removed 749M Anna's Archive URLs from Its Search Results
#75Earlier quoted context omitted.
They’re… yes. Yes, that’s exactly what they have done and continue to do. Are you familiar with it?
I think the comment is saying Google was also doing that.
Re: Google Removed 749M Anna's Archive URLs from Its Search Results
#76Earlier quoted context omitted.
Google hides the most relevant results on the 3rd page. It was confirmed in trial disclosures a few months ago. Their concern isn’t public search.
Google only ever returns a maximum of Bing only returns 900. Kagi only 200. Deep search and surfing is pretty much gone on all major search "engines".
What's not cool about Google is that now it hits you with AI slop with dubious quality right at the top, followed by a page of sponsored results, followed by some potentially useful results, followed by an entire ocean of spam traps and clone sites and really shady results with exotic never-seen-before TLDs that leaves you wondering whether clicking on a link will get in a hostile database. That's what's not cool about Google: is that you can't use it to search the web anymore.
Re: Google Removed 749M Anna's Archive URLs from Its Search Results
#77Earlier quoted context omitted.
1. Your chatbot doesn't have its own internet scale search index. 2. You're being given information that may or may not be coming in part from junk sites. All you've done is give up the agency to look at sources and decide for yourself which ones are legitimate.
I’m quite happy trading off the agency of wading through trash to an LLM. In fact, I would say that’s something they’re pretty good at.
Re: Google Removed 749M Anna's Archive URLs from Its Search Results
#78Feels weird to say but I have found using Yandex of all places an excellent search engine for content that get taken down by DMCA requests. Eg if you want to watch a movie that's not on Netflix using a web stream the search results are far better. Feels like Google circa 2005.
I've been playing around with a variety of search engines such as Kagi, Startpage, Ecosia, DDG. All of them are better than google in finding relevant results. Lol Google is way too "personalized".
Re: Google Removed 749M Anna's Archive URLs from Its Search Results
#79Re: Google Removed 749M Anna's Archive URLs from Its Search Results
#80Man I need to get around to downloading the z-archive torrents before annas archive is taken down. If I eliminate large PDFs and non english books I think I can fit it on two 32 TB drives with BTRFS z-std compression max setting. https://annas-archive.org/torrents
How large? Isn't that going to result in an arbitrary filter of books? In other domains, large PDFs are due to PDF production errors, such as using color or needlessly high resolution, and not so much due to the volume of content - at least for text.