Live data from Hacker News

A search engine that favors text-heavy sites and punishes modern web design

search.marginalia.nu

251–260 of 735 posts

Re: A search engine that favors text-heavy sites and punishes modern web design

#252
post #7

Where does the data come from? Do you index the whole web yourself? I see it totally impossible for a personal project. I'm very curious about that.

I do indeed index the web myself. Not the entire web, just a subset of it. The crawler quickly loses interest in javascript:y websites and only indexes at depth those websites that are simple. It also focuses on websites in English, Swedish and Latin and tries to identify and ignore the rest (best-effort). You'd be surprised how much you can do with modern hardware if you are scrappy. The current index is about 17.7…

Cool, I've been thinking on this topic a bit lately. Crawling is indeed not that hard of a problem. Google could do it 23 years ago. The web is a bit bigger now of course but it's not that bad. Those numbers are well within the range of a very modest search cluster (pick your favorite technology; it shouldn't be challenging for any of them). 10x or 1000x would not matter a lot for this. Although it would raise your cost a little.

The hard problem is indeed separating the good stuff from the bad stuff; or rather labeling the stuff such that you can tell the difference at query time. Page rank was nice back in the day; until people figured out how to game things. And now we have bot farms filling the web with nonsense to drive political agendas, create memes, or to drown out criticism. Page rank is still a useful ranking signal; just not by it self.

The one thing no search engine has yet figured out is reputability of sources. Content isn't anonymous mostly. It's produced and consumed by people. And those people have reputations. Bot content is bad because it comes from sources without a credible reputation. Reputations are built over time and people value having them. What if we could value people's appreciation relative to their reputability? That could filter out a lot of nonsense. A simple like button + a flag button combined with verified domain ownership (ssl certificates) could do the trick. You like a lot of content that other people disliked, your reputation goes down the drain. If you produce a lot of content that people like, your reputation goes up. If a lot of reputable people flag your content, your reputation tanks.

The hard part is keeping the system fair and balanced. And reputability is of course a subjective notion and there is a danger of creating recommendation bubbles, politicizing certain topics, or even creating alternative reality type bubbles. It's basically what's happening. But it's mostly powered by search engines and social media that actually completely ignore reputability.

Re: A search engine that favors text-heavy sites and punishes modern web design

#255

Earlier quoted context omitted.

This sort of optimization is why simple recipes are typically found at the end of a rambling pointless blog post now. Still, the best way to break SEO is to have actual competition in the search space. As long as SEO remains focused on Google there is an opportunity for these companies to thrive by evading SEO braindamage.

That sort of recipe blog hasn't happened just for SEO. It's also a bit of a "two audiences" problem: if you are coming to that food blogger from a search you certainly would prefer the recipe first and then maybe any commentary on it below if the recipe looks good. If you are a regular reader of that food blogger you are probably invested in the stories up top and that parasocial connection and the recipes themselves…

I see your point, but argue you've misidentified the two audiences.

One audience matches your description and is the invested reader. They want that blogger's story telling. they might make the recipe, but they're a dedicated reader.

The other audience is not the recipe-searcher, but instead Google. Food bloggers know that recipe-searchers are there to drop in, get an ingredient list, and move on. They won't even remember the blog's name. So the site isn't optimized for them. It's optimized for Google.

"Slow the parasitic recipe-searcher down. They're leeches, here for a freebie. Well they'll pay me in Google Rank time blocks."

Re: A search engine that favors text-heavy sites and punishes modern web design

#256
post #220

There's probably a more suitable term than "modern" that we should generally be using, since "modern" consistently has a positive connotation.

Dunno, I prefer to use as neutral or positive terminology even when I talk about things I don't like. I think it very easily comes off as juvenile ranting when you start throwing around terms with strong negative connotations.

Re: A search engine that favors text-heavy sites and punishes modern web design

#257

Earlier quoted context omitted.

As long as few people use it, it will be great. Rest assured that the moment it becomes popular, the people who want to game it will appear.

> the people who want to game it will appear. So just add human review to the mix, if a site is obviously trying to game the system (listicles, seo spam etc) just drop and ban them from the search index.

Congratulations, you've just invented negative SEO.

Re: A search engine that favors text-heavy sites and punishes modern web design

#258
post #92

Earlier quoted context omitted.

Perhaps a counter example, something that is interesting. Anecdotally. This, of all things, is the top result in my search: https://tft.brainiac.com/archive/0303/msg00037.html . Which is strange to me because I don't recognize tft.brainiac. I click, it's a list of biological relationships among Hymenoptera, including a reference to genus of the wasps I studied, presumably in a biological relationship (host/parasite)…

Yes, thanks. That helps. If I understand it correctly, you're interested in bits and pieces of new information that's indirectly related to your object of interest. Degree 2 and 3 in Six Degrees of Kevin Bacon, so to speak. You know degree 0 like the back of your hand and you've seen almost everything closely connected. Finding novel, interesting things is getting more difficult. Have you thought about cataloging all…

Exactly.

> Have you thought about cataloging all the related stuff you stumble upon? Something in between loose notes and what Moby Dick is to cetology.

Tongue in cheek- new app time, to facilitate this. It should have the name "Degree4". Entries can only be made if degrees 2 and 3 are "defined". Scoffs at degrees 5 and 6, just because. Startup developing can probably unethically seed content by mining https://www.everything2.com/. Should use concepts of "AI" and "persistent homology"... profit!

But no, I don't outside a mental note. Closest I would come would be adding '!! ' to my potwiki text notes (see my past comments) if its something I want to have come back with a grep, or think might be interesting to explore "when I retire". If it's a scientific fact in my field after researching it further it would go into this https://taxonworks.org (or its precursor).

Re: A search engine that favors text-heavy sites and punishes modern web design

#259
This has been needed in my life for a while. I am growing really apathetic about the internet lately, but I realize that is because my entry point is always a google search.

I miss finding blog posts and scholarly articles in long form. I hate the SEO sites with unreadable UI because the information in them is often a lot lower quality as well.

Re: A search engine that favors text-heavy sites and punishes modern web design

#260
Pretty cool. I am not sure yet how useful, but cool it is.

However, it seems that it currently does not support non-Latin alphabets. Which I understand in an early version. Still, it's handling of such "exception cases" could be improved:

when I search for a Russian word, say "Аквариум", I get >, which is rather rude...

Post reply on HN