Live data from Hacker News

Alexandria Search

alexandria.org

51–60 of 181 posts

Re: Alexandria Search

#51
I love the shortcut Alexandria takes by indexing Common Crawl instead of crawling the web themselves. It's how I would have bootstrapped a new search engine. In a future iteration they can start crawling themselves, if there is sufficient interest from the public.

Searching is screamingly fast.

The index seems stale, though. Alexandria, how old is your index?

How long did it take you to create your current index? Is that your bottleneck, perhaps, that it takes you a long time (and lots of money?) to create a Common Crawl index?

Re: Alexandria Search

#52
For my first search of "GFlowNetworks" (which the search bar suggested) It said: Found 5,887 (or something) results, but showed no results

For my second I searched my name and got a Wikipedia article about a show I've never heard of which didn't have my name anywhere in it.

For my third I searched "GFlowNetworks" again and it said Found 2,656,844 results in 1.61s, but showed no results again

Re: Alexandria Search

#53
post #51

I love the shortcut Alexandria takes by indexing Common Crawl instead of crawling the web themselves. It's how I would have bootstrapped a new search engine. In a future iteration they can start crawling themselves, if there is sufficient interest from the public. Searching is screamingly fast. The index seems stale, though. Alexandria, how old is your index? How long did it take you to create your current index? Is…

> The index seems stale, though. Alexandria, how old is your index?

Common crawl indexes about once every 40 days, the current crawl's data is through January 2022, so it's 1.5 months old at best.

Re: Alexandria Search

#55

There have been a few search engines out recently. I'm curious how people evaluate them quickly. I've realized my searching is basically optimized for google and the web that has grown up around it. Also, in 1998 I wasn't as aware of what was out there as I am now. It's pretty rare (even if its possible) that I do a search and come across a completely new site that I haven't heard of before, for anything nontrivial.…

> There have been a few search engines out recently

I'd like to try them out, could you mention which?

Re: Alexandria Search

#56
post #55

There have been a few search engines out recently. I'm curious how people evaluate them quickly. I've realized my searching is basically optimized for google and the web that has grown up around it. Also, in 1998 I wasn't as aware of what was out there as I am now. It's pretty rare (even if its possible) that I do a search and come across a completely new site that I haven't heard of before, for anything nontrivial.…

> There have been a few search engines out recently I'd like to try them out, could you mention which?

you.com and kagi.com off the top of my head

Re: Alexandria Search

#57
Really like the minimal UI and the speed! Great work.

A few of my test searches came up with very useful results. However, one disappointment was searching for a javascript function, for example "javascript array splice", and the MDN site was not in the results. Adding "MDN" or "Mozilla" to the search did not help either.

Re: Alexandria Search

#58
post #15
post #9

I think the fact that after a long while there are new search engines (Kagi was introduced very recently on HN, now this) should be a wake up call for Google - their search has lost some shine for quite a while. Hopefully something will come out of this - competition is good.

At some point, AI and NLP and raw processing power will have progressed so much that "search" is not a problem anymore, and I think we're getting there. Google can up their game but it won't matter much. The only thing they have left is brand recognition.

IMO search has had its goalpost moved. It used to be about scale, technical challenges, bandwidth, storage, etc. It is still about that, but a significantly harder challenge to solve has come up: searching in a malicious environment. SEO crap nowadays completely dominates search, Google has lost the war.

Simply put, I believe that Google sucks at search, in the modern context. It is great at indexing, it has solved phenomenal technical challenges, but search it has not solved. Why do I have to write site:stackoverflow.com or site:reddit.com to skip the crap and go to actual content? Why can my brain detect blogspam garbage in 0.5 seconds of looking but billion dollar company Google will happily recommend it as the most relevant result above a legitimate website?

I feel this 12 year old XKCD is still relevant: https://xkcd.com/810/ .

Post reply on HN