Live data from Hacker News

A search engine that favors text-heavy sites and punishes modern web design

search.marginalia.nu

231–240 of 735 posts

Re: A search engine that favors text-heavy sites and punishes modern web design

#231

You should monetise this with amazon affiliate links that are relevant to each search. And then use that money to keep this project going. Google is fantastic, but it has become something different from what it was, the company and the product. It is so refreshing to see a modern tool that encourages exploration of the actual world wide web.

I might add a donate button or something if people want to help support the project, hardware isn't cheap and all. But I have a job and decent income. I think if this search engine became the way I earned money, it would influence the project in a bad way, and corrupt its purpose, which is to help people explore the less-commercial internet.

Appreciated. The more things fill up with monetizing shit, the more I stay away. There's something beautiful in having higher purposes than grubbing for cash.

Re: A search engine that favors text-heavy sites and punishes modern web design

#233
post #7

Where does the data come from? Do you index the whole web yourself? I see it totally impossible for a personal project. I'm very curious about that.

I do indeed index the web myself. Not the entire web, just a subset of it. The crawler quickly loses interest in javascript:y websites and only indexes at depth those websites that are simple. It also focuses on websites in English, Swedish and Latin and tries to identify and ignore the rest (best-effort). You'd be surprised how much you can do with modern hardware if you are scrappy. The current index is about 17.7…

> It also focuses on websites in English, Swedish and Latin and tries to identify and ignore the rest

When I search for Japanese terms, it "says needs to be a word", which wasn't the best error message. Maybe the error message should say something like "sorry, your language isn't support yet"?

Re: A search engine that favors text-heavy sites and punishes modern web design

#235
post #175

Earlier quoted context omitted.

It would be nice if we could pipe search engines.

Definitely; We could create a meta search engine that queries them all, in desktop application format. Let's name it after a famous old scientist, and maybe add the year to prove it's modern: Galileo 2021.

I need that with a simpler interface, so I call it after a famous dedective: Sherlock.

Re: A search engine that favors text-heavy sites and punishes modern web design

#236
post #7

Where does the data come from? Do you index the whole web yourself? I see it totally impossible for a personal project. I'm very curious about that.

I do indeed index the web myself. Not the entire web, just a subset of it. The crawler quickly loses interest in javascript:y websites and only indexes at depth those websites that are simple. It also focuses on websites in English, Swedish and Latin and tries to identify and ignore the rest (best-effort). You'd be surprised how much you can do with modern hardware if you are scrappy. The current index is about 17.7…

Do you know what proportion of the texty web instructs unknown crawlers to go away (or blocks them)?

Re: A search engine that favors text-heavy sites and punishes modern web design

#237

I wish we could configure Google's algorithm to our needs, and blacklist websites.

It could get tedious depending on how many sites you want to block, but you can add "-site:google.com" to exclude google.com, for instance.

I mean a blacklist system like Twitter's, where you block a website forever. Pinterest would be the first to go.

Re: A search engine that favors text-heavy sites and punishes modern web design

#238
Ivermectin (marginalia): https://search.marginalia.nu/search?query=ivermectin+

Ivermectin (Google): https://www.google.com/search?q=ivermectin

The difference in the overall _thrust_ of the results is remarkable.

Very interesting! Thanks for building it.

Re: A search engine that favors text-heavy sites and punishes modern web design

#239
post #175

Earlier quoted context omitted.

It would be nice if we could pipe search engines.

Definitely; We could create a meta search engine that queries them all, in desktop application format. Let's name it after a famous old scientist, and maybe add the year to prove it's modern: Galileo 2021.

Not an app, but probably comes quite close in all other respects: https://metager.org

Re: A search engine that favors text-heavy sites and punishes modern web design

#240

Earlier quoted context omitted.

This search engine pretty much takes everything that Google is doing and does the opposite. For instance, Google has decided that "relevant" usually also means "recent". Thus, when searching for something on Google, you mainly get results from blogspam farms and almost never do you see anything more than a few years old. An implication of this is that old sites tend to disappear (either into obscurity or by being tak…

What "fundamental redeeming quality" about uninstalling AIM from Windows 3.x motivated making that the 3rd result for "dog"? The 5th result is a tutorial on CSS. This search engine decided it's relevant because it has "dog" in the URL. Is that a better reasoning than Google's? https://htmldog.com/guides/css/beginner/ Core Web Vitals ranks sites higher that perform well. Text-heavy sites that are also optimized and re…

What are you searching for when you enter the query "dog", keeping in mind the search engine deliberately does not examine synonyms or and deliberately seeks out the path less taken?

Dog facts? Then search "dog facts"

Famous dogs? Then search "famous dogs"

Rappers? Try "snoop dogg"

Post reply on HN