Live data from Hacker News

Alexandria Search

alexandria.org

121–130 of 181 posts

Re: Alexandria Search

#121
This is really lovely. I think the search results are useful, if you're looking for more static information. Very pleasant, to see only results.

I think, I found a minor bug, while (of course) searching for my homepage:

https://www.alexandria.org/?c=&r=&q=www.hadjian.com

The status line below the search box says, that it found many results, but the results are empty. Also, when hitting F5 a couple of times, the number jumps around.

Keep up the great work. I think there is a lot potential to Common Crawl and things built on top of it.

Re: Alexandria Search

#122

There have been a few search engines out recently. I'm curious how people evaluate them quickly. I've realized my searching is basically optimized for google and the web that has grown up around it. Also, in 1998 I wasn't as aware of what was out there as I am now. It's pretty rare (even if its possible) that I do a search and come across a completely new site that I haven't heard of before, for anything nontrivial.…

> I'm curious how people evaluate them quickly.

To paint with a broad brush, I look at three criteria:

1. Infoboxes ("instant answers") should focus on site previews rather than trying to intelligently answer my question. Most DuckDuckGo infoboxes are good examples of this; Bing and Google ones are too "clever".

2. Organic results should be unique; most engines are powered by a commercial Bing API or use Google Custom Search. Compare results with a Bing or Google proxy (duckduckgo, startpage, etc) to avoid personalized results. Monitor queries over time to see if SERPs change in ways that diverge from Google/Bing/Yandex.

3. "other" stuff. Common features I find appealing include area-specific search (Kagi has a "non-commercial lens" mostly powered by its Teclis index; Brave is rolling out "goggles"), displaying additional info about each result (Marginalia and Kagi highlight results with heavy JS or tracking), user-driven SERP personalization (Neeva and Kagi allow promoting/demoting domains), etc.

And always check privacy policies, TOS, GDPR/CCPA compliance, etc.

> Google is now almost a convenience. If I have a coding question, I search for "turn list of tensors into tensor" or whatever but I'm really looking for SO or the pytorch documentation, and I'll ignore the geeksforgeeks and other seo spam that finds it's way it. It's almost like google is a statistical "portal" page,

I like engines like Neeva and Kagi that allow customizing SERPs by demoting irrelevant results; I demote crap like GFG, w3schools, tutorialspoint, dev(.)to, etc. and promote official documentation. Alternatively, you can use an adblocker to block results matching a pattern: https://reddit.com/hgqi5o

Re: Alexandria Search

#123

There have been a few search engines out recently. I'm curious how people evaluate them quickly. I've realized my searching is basically optimized for google and the web that has grown up around it. Also, in 1998 I wasn't as aware of what was out there as I am now. It's pretty rare (even if its possible) that I do a search and come across a completely new site that I haven't heard of before, for anything nontrivial.…

I'm surprised you are not a fan of geeksforgeeks. While each of their webpages have substantially less content than the pytorch docs or SO result, I find that they get to the point instantly. My mean time to solution from G4G is definitely smaller than SO.

I generally find that sites like SO, GFG, etc. often play the role of "Reading the Docs as a Service". I prefer using them only after official documentation or specifications fail me. When I want an opinionated answer, I just ping some people I already know or check the personal websites of the developers of the language/tool I'm using. If I have further questions, I check the relevant IRC channel. Sites like SO are a "last resort" for me.

In other words, I'd rather see these at the bottom of the SERP than the top, but I wouldn't want to completely eliminate them.

Re: Alexandria Search

#124
post #43

Is there a web that pools multiple search engines results?

SearX and Searxng are the most common options, but instances often get blocked by the engines they use. Users need to switch between instances quite often.

eTools.ch uses commercial APIs so it doesn't get blocked, but it might block you instead (very sensitive bot detection).

Dogpile is one of the older metasearch engines, but I think it only uses Bing- and Google-powered engines.

Re: Alexandria Search

#126

This actually makes me want to build my own web crawler and search

Founder here, I suggest you start by not implementing a crawler but use commoncrawl.org instead. The problem with starting a web crawler is you will need a lot of money and almost all big websites are behind cloudflare so you will be blocked pretty quickly. Crawling is a big issue and most of the issues are non-technical.

I've heard from other people who run engines (Right Dao, Gigablast) that this is a major problem; Common Crawl does look helpful, but it's not continuously updated. FWIW, Right Dao uses Wikipedia as a starting point for crawling. Kiwix makes pre-indexed dumps of Wikipedia, StackExchange, and other sites available.

Some sort of partnership between crawlers could go a long way. Have you considered contributing content back towards the Common Crawl?

Re: Alexandria Search

#127

There have been a few search engines out recently. I'm curious how people evaluate them quickly. I've realized my searching is basically optimized for google and the web that has grown up around it. Also, in 1998 I wasn't as aware of what was out there as I am now. It's pretty rare (even if its possible) that I do a search and come across a completely new site that I haven't heard of before, for anything nontrivial.…

> I'm curious how people evaluate them quickly. Speaking about myself; I cold turkey migrated to DDG ~2 months ago. So far I've had to resort to Google search 10 times or so. One thing I miss though is Google's nice visualisation of fast changing results e.g., match scores. For example: https://imgur.com/a/Q5nZkjo

DDG's organic link results are from Bing, sans personalization. DuckDuckGo advertises using "over 400 sources", which means that at least 399 sources only power infoboxes ("instant answers") and non-generalist search, such as the Video search.

Re: Alexandria Search

#128

There have been a few search engines out recently. I'm curious how people evaluate them quickly. I've realized my searching is basically optimized for google and the web that has grown up around it. Also, in 1998 I wasn't as aware of what was out there as I am now. It's pretty rare (even if its possible) that I do a search and come across a completely new site that I haven't heard of before, for anything nontrivial.…

maybe the solution is to make google itself reddit style. let users downvote the seo spam websites and allow them to be downranked. sure it opens the door for a different kind of abuse... but maybe that problem is more fixable?

Re: Alexandria Search

#129

There have been a few search engines out recently. I'm curious how people evaluate them quickly. I've realized my searching is basically optimized for google and the web that has grown up around it. Also, in 1998 I wasn't as aware of what was out there as I am now. It's pretty rare (even if its possible) that I do a search and come across a completely new site that I haven't heard of before, for anything nontrivial.…

> I've realized my searching is basically optimized for google Is it just me, or I feel like Google does not provide anymore good results for me. Like every time I search something completely out of my knowledge, like "How to purchase a property in Mexico", it will give me 100+ results of some results with autogenerated content like "10 best places to buy property in Mexico". And the only way to fix that would be to…

> Is it just me, or I feel like Google does not provide anymore good results for me.

I am starting to suspect that there might be nothing to find.

I just don't think people (other then the tech oriented) are creating websites and running forums - and why would they? Reddit might be be only place you _can_ find that type of content. What should search engines do then?

With a tiny number of exceptions, it might be that people chat on reddit, read Wikipedia, ask questions on the stackexchange network/Quora, local communities use facebook groups, and businesses have a wordpress site with nothing more then a bit of fluff, a phone number and an email address.

Re: Alexandria Search

#130
post #55

There have been a few search engines out recently. I'm curious how people evaluate them quickly. I've realized my searching is basically optimized for google and the web that has grown up around it. Also, in 1998 I wasn't as aware of what was out there as I am now. It's pretty rare (even if its possible) that I do a search and come across a completely new site that I haven't heard of before, for anything nontrivial.…

> There have been a few search engines out recently I'd like to try them out, could you mention which?

I listed a bunch over at https://seirdy.one/2021/03/10/search-engines-with-own-indexe..., and I'm always adding more.

I first discovered Alexandria in early February: https://git.sr.ht/~seirdy/seirdy.one/commit/935b55f10f9024ee...

Around the same time, I also discovered sengine.info, Artado, Entfer, and Siik. By sheer coincidence they all were mentioned to me or decided to crawl my site within the same couple weeks. So yes, from my perspective there have been more than a few new smaller engines getting active on the heels of bigger names like Neeva, Kagi, Brave Search, etc.

Post reply on HN