Live data from Hacker News

A search engine that favors text-heavy sites and punishes modern web design

search.marginalia.nu

121–130 of 735 posts

Re: A search engine that favors text-heavy sites and punishes modern web design

#122
This is a fascinating tool, I estimated that the corpus of the factual web was between 1 and 10 TB when I last played around with BigQuery using domain names which had low amounts of click bait. Seeing these search results I suspect my estimate was off by a couple orders of magnitude.

Although a search for "Fractional Reserve Banking" shows that some further ranking improvements can be made to exclude unrelated results, and potentially penalize old conspiracy sites.

https://search.marginalia.nu/search?query=fractional+reserve...

Re: A search engine that favors text-heavy sites and punishes modern web design

#123
I love it. Even though it didn't give me the results I was looking for. I searched "new york fishing license", and it didn't give me any links to the actual new york fishing license websites. But it did give me a ton of really cute little websites related to lakes and fishing in New York. This one has amazing information about fishing all over Western New York: http://www.huntfishnyoutdoors.com/fishing.php

Re: A search engine that favors text-heavy sites and punishes modern web design

#125
post #50

Fascinating. I studied an "obscure" group of insects. My go-to search term to test an engine is their family name as it is a rarely used word and I know most (all?) of the major data sources that have accumulated data on it. When Wolfram Alpha added species names, I checked with the name, boring, Duck Duck, boring, Google (well we know Google isn't for search anymore, it's absolutely horrible) boring, Bing, boring...…

I'm intrigued by this experiment but I can't visualize it. What do you mean by boring results? Would combing through a library (the one with paper books) also produce boring results? What's your ideal results?

In part, by boring results I mean I instantly recognize the top results, and I know exactly what will be in them, and I know which ones will actually contain potentially interesting new stuff, i.e. _I didn't have to search for these, I'd go their directly_. Then next results are all obscure, and I've already visited them, and/or I know they are historical and not something I have to revisit.

With this engine with at least 1/2 the links (to be fair there were I suppose the magic in this engine would have to be alerting the searcher that they found more of this type of link, as once I visited the 10 or so sites they would fall back into the "been there, done that" link category that Google appends somewhere after the ads and "big" sites, mixed in with a million search term spam sites, etc.

Re: A search engine that favors text-heavy sites and punishes modern web design

#126

from About page: > If you search for "Plato", you might for example end up at the Canterbury Tales. Go looking for the Canterbury Tales, and you may stumble upon Neil Gaiman's blog. I know it is just a suggestion, but had to try searching both, with no luck in getting the expected unexpected.

Yeah I did some work very recently aimed at improving the relevance a bit. It was a bit too random in the state it was before. Now it, perhaps, isn't random enough anymore.

It looks very nice anyway, great job! I did try with other queries and results were in general interesting.

Re: A search engine that favors text-heavy sites and punishes modern web design

#127
post #68

Earlier quoted context omitted.

I'm intrigued by this experiment but I can't visualize it. What do you mean by boring results? Would combing through a library (the one with paper books) also produce boring results? What's your ideal results?

There’s certain grey literature that’s not captured in university library federated searches nor easily found with mainstream search engines.

There are decades of academic research not digitized. The digitization window used to only hit around 1990, I haven't looked at it hard recently, but I suspect this still remains true for many important journals. This is grey only to those who do not know how to use a library.

Re: A search engine that favors text-heavy sites and punishes modern web design

#128

I like it. Coincidentally, the other day I was daydreaming about a search engine that favors sites that are updated less frequently. The thought being, the kinds of labors of love that characterized the 1990s Web that I still sometimes miss are still out there, it's just harder to find them amidst the flood of SEO dreck. So perhaps they could be made discoverable again with the help of a contrarian search engine that…

I had this problem recently trying to fix an Atari. There's a guy out there who has ton's of guides on doing video out mods but newer guide references the older. However googling the OG guide didn't find it so I manually scoured his old web page.

Re: A search engine that favors text-heavy sites and punishes modern web design

#130
> New: You can now look up dictionary definitions for words. If you for example don't know what the definition of is is, you can inquire thus: define:is.

Oh man, I love subtle jabs and tongue in cheek writing like this. Very Robin Williams-esque.

Post reply on HN