Live data from Hacker News

Ask HN: Is there a search engine which excludes the world's biggest websites?

news.ycombinator.com

131–140 of 236 posts

Re: Ask HN: Is there a search engine which excludes the world's biggest websites?

#132
post #4

This is a great question, I also want a way to search the internet but exclude all major media domains as well as any company over a certain size. So I just want to search through old blogs, SO, non-corporate social media, weird forums, etc. There are so many cool things I remember reading on the web like 10-20 years ago that still exist that are so buried now on Google they might as well not exist. Nowadays searchin…

This is somewhat ironic because 20 years ago, hobbyists would frequently put their obscure personal pages on Geocities and other large corporation's web space.

20 years ago, it was extremely common for your ISP to give you 5MB or whatever of space to use. users.ispname.com/~yourname or whatever. It was great, tbh, since anyone with cuteftp and notepad could publish to the world.

Re: Ask HN: Is there a search engine which excludes the world's biggest websites?

#133
Reminds me of this classic pg essay: http://www.paulgraham.com/ambitious.html

Specifically this quote: "The way to win here is to build the search engine all the hackers use. A search engine whose users consisted of the top 10,000 hackers and no one else would be in a very powerful position despite its small size, just as Google was when it was that search engine."

There has been a lot of grumblings about the state of search these days. Maybe the time is nigh for a new search engine?

Re: Ask HN: Is there a search engine which excludes the world's biggest websites?

#134
post #130

Earlier quoted context omitted.

Excuse the incivility, but no. COVID-19: disease caused by SARS-CoV-2 SARS-CoV-2: strain of SARS-CoV SARS-CoV: severe accute respiratory syndrome coronavirus Coronavirus: virus that causes respiratory diseases in mammals, such as SARS (SARS-CoV) MERS (MERS-CoV), and COVID-19 (SARS-CoV-2)

>SARS-CoV-2: strain of SARS-CoV Excuse the incivility, but no. SARS-CoV-2 is not a strain or type of SARS-CoV. The viruses share ancestors, but SARS-CoV-2 did not come directly from SARS-CoV. SARS-CoV and SARS-CoV-2 are in the category of beta coronaviruses[0]. "The whole genome-based phylogenetic analysis presented that two Bat SARS-like CoVs (ZXC21 and ZC45) were the closest relatives of SARS-CoV-2."[1] [0] https:/…

Excuse the incivility once again, but no.

While we're on the topic of linguistic pedantary, strain isn't exclusive to direct mutations from a parent genome. Strains, like much of biological taxonomy, are a human abstraction to make communication of the idea of -- in this case -- "a virus sharing similar properties to coronaviruses that cause severe acute respiratory syndrome" -- albeit this is a very simplified definition for the sake of brevity.

SARS is caused by SARS-CoV-1 and COVID-19 is caused by SARS-CoV-2.

Rather, if we would like to be absolutely correct about these classifications, we would say SARS-CoV-1 and SARS-CoV-2 are both strains of SARSr-CoV (Severe accute respiratory syndrome related coronavirus), which in itself is a species, an abstract concept used to group related organisms into a convenient umbrella term.

There is no "eukaryote" organism the same way there is no "SARSr-CoV" organism. The added "r" was a recent addition when COVID-19 was discovered.

I will cede that I didn't specify this last point, and you were correct to point it out.

Re: Ask HN: Is there a search engine which excludes the world's biggest websites?

#135
For

> Ask HN: Is there a search engine which excludes the world's biggest websites?

> Discovering unknown paths of the web seems almost impossible with google et al..

> Are there any earch engines which exclude or at least penalize results from, say, top 500 websites?

Let's back up a little and then try for an answer:

Some points:

(1) For some qualitative exclamation, there is a LOT of content on the Internet.

(2) There are in principle and no doubt so far significantly in practice a LOT of searches people want to do. The search in the OP is an example.

(3) Much like in an old library card catalog subject index, the most popular search engines are based heavily on key words and then whatever else, e.g., page rank, date, etc.

So: (1) -- (3) represent some challenges so far not very well met: In particular, we can't expect that the key words, etc. of (3) will do very well on all or nearly all the searches in (2) for much of the content in (1).

And the search in the OP is an example of a challenge so far not well met.

Moreover, the search in the OP is no doubt just one of many searches with challenges so far not well met.

Long ago, Dad had a friend who worked at Battelle, and IIRC they did a review of information retrieval that concluded that keyword search covers only a fraction, maybe ballpark only 1/3rd, of the need for effective searching. And the search in the OP is an example of what is not covered because the library card catalog did not index size of the book or Web site! :-)!

Seeing this situation, my rough, ballpark estimate has been that the currently popular Internet search engines do well on only about 1/3rd of the content on the Internet, searches people want to do, and results they want to find.

So, I decided to see what could be done for the other 2/3rds.

I started with some not very well known or appreciated advanced pure math; it looks like useless, generalized abstract nonsense, but if calm down, stare at it, think about it, ..., can see a path for a solution. Although I never thought about the search in the OP until now, in principle the solution should work also for that search. Or, the math is a bit abstract and general which can translate in practice to doing well on something as varied as the 2/3rds.

Then for the computing, I did some original applied math research.

Using TeX, I wrote it all up with theorems and proofs.

So, the project is to be a Web site. While in my career I've been programming for decades, this was my first Web site. I selected Windows and .NET, and typed in 100,000 lines of text with 24,000 statements in Visual Basic .NET (apparently equivalent in semantics to C# but with syntactic sugar I prefer).

The software appears to run as intended and well enough for significant production.

I was slowed down by one interruption after another, none related to the work.

But, roughly, ballpark, the Web site should be good, or by a lot the best so far, for the 2/3rds and in particular for the search in the OP.

So, for

> Ask HN: Is there a search engine which excludes the world's biggest websites?

there's one coded and running and on the way to going live!

I intend to announce an alpha test here at HN.

Re: Ask HN: Is there a search engine which excludes the world's biggest websites?

#136
My hobby project is https://random.surf (works better on desktop than mobile).

I share that same desire to visit the web less travelled. I want to discover interesting sites that deserve to be bookmarked because they will never show up in a search engine.

Re: Ask HN: Is there a search engine which excludes the world's biggest websites?

#137
post #130

Earlier quoted context omitted.

>SARS-CoV-2: strain of SARS-CoV Excuse the incivility, but no. SARS-CoV-2 is not a strain or type of SARS-CoV. The viruses share ancestors, but SARS-CoV-2 did not come directly from SARS-CoV. SARS-CoV and SARS-CoV-2 are in the category of beta coronaviruses[0]. "The whole genome-based phylogenetic analysis presented that two Bat SARS-like CoVs (ZXC21 and ZC45) were the closest relatives of SARS-CoV-2."[1] [0] https:/…

Excuse the incivility once again, but no. While we're on the topic of linguistic pedantary, strain isn't exclusive to direct mutations from a parent genome. Strains, like much of biological taxonomy, are a human abstraction to make communication of the idea of -- in this case -- "a virus sharing similar properties to coronaviruses that cause severe acute respiratory syndrome" -- albeit this is a very simplified defin…

>we would say SARS-CoV-1 and SARS-CoV-2 are both strains of SARSr-CoV

Thank you for making my point, again.

Re: Ask HN: Is there a search engine which excludes the world's biggest websites?

#138
post #24

Earlier quoted context omitted.

It is a carefully curated directory, which is problematic. For example, I submitted Pizza Hut's archived original web page [1], but it wasn't added. Even for a search engine exposing niches, updating a directory manually will likely be too slow, unless the directory is maintaining a single nich (e.g., unladen airspeed of every species of swallow), but then we end up with some insane number of search engines and how t…

Especially if you’re focussing on evergreen information, there’s no reason why people can’t have their own personalized crawler and index— I’ve occasionally thought about rolling my own with a browser extension that lets me add seeds at the click of a button.

I've been working on something like this for my own use - I'm not a fan of browser-based history. My home-rolled solution is starting to be good enough where I can use it to easily find exactly what I'm looking for, assuming I've previously read it, by both searching the title and URL, as well as the content on that page (my major gripe with "History" in Chrome and Firefox is that it doesn't search the page content, and if it did, syncing it would have major privacy concerns).

The problem I'm running into is that I still have to use major search engines to find new content, way more than I'd like. I hope to make my local service available open source once I have 'federated' history search working, so that we can have a primitive search engine and share with people we trust. Also need to work out some security issues - it's scary having all the content you read and see on your home network, protected only by your hackily-patched-together security.

EDIT: Actually I'd like to elaborate a bit more in case anybody actually reads this and has any ideas. On the desktop side, it's pretty easy. Initially started out MITMing my own traffic with a self-signed cert added as a root cert to all my machines. This only works on my home network, so I did a VPN thing. This was way to clunky and the security concerns are innumerable. I ended up biting the bullet and writing a chrome extension which works wonderfully, except for some slight performance issues.

However, I wish to also archive my phone content - I read just as much on my phone as my computer. I can do it on Android with the MITM process, but the same issues as above still apply, and it doesn't work with iOS (at least I can't find a way).

I'm thinking of taking an open source project, like Firefox/Fennec and building it in to the app itself. In that case it may make sense to forgo the browser extension and just roll my own forked browser on every platform, even iOS. I don't know much about iOS dev though.

Re: Ask HN: Is there a search engine which excludes the world's biggest websites?

#140

Simply removing Pinterest would be a huge step in the right direction.

I use an add-on called Unpinterested! to remove Pinterest results from my Google search results:

https://github.com/sellomkantjwa/unpinterested

All it does is add -site:pinterest.com to the search bar for image results (can be configured to also do it for Web results), but it gets the job done.

Post reply on HN