Live data from Hacker News

A search engine that favors text-heavy sites and punishes modern web design

search.marginalia.nu

471–480 of 735 posts

Re: A search engine that favors text-heavy sites and punishes modern web design

#471

Earlier quoted context omitted.

Imagine if you were looking for the movie.

I tend to prefer Wikipedia for movies. The exception is actor headshots if I'm trying to identify someone, which Wikipedia lacks for licensing reasons, but otherwise Wikipedia tends to be better than IMDB for most needs. Wikipedia has an IMDB link on every article anyway. Another need I guess might be reviews, for which RT or MC are better than IMDB: not sure if either of those two will fare better than IMDB in this…

I find IMDb to be more convenient than RT/MC/Wikipedia for finding release dates of movies - nearly every other website lists only the American release date, maybe one or two others if the movie was disproportionately popular in certain regions.

Re: A search engine that favors text-heavy sites and punishes modern web design

#472
As a quick test, I searched for the name of one of my favorite game series: "Baldur's Gate" (on its own, no qualifiers, properly spelled - I would usually spell it "baldurs gate" on Google, but I decided to give this one the best chance). I search for info around video games a lot, so that's quite representative of a good chunk of my web searches, and I pretty much know the top sites Google would give me for that query (on its own, without any further qualifiers).

The results were all either barely relevant, outdated (sites that covered the game back in the 90s/2000s before it was re-released), at best tangentially relevant or complete garbage noise. Some of the most highly relevant pages (such as the Steam store listing, the fandom wiki, the publisher/developer's forums for the re-releases, the Baldur's Gate 3 website and the subreddit) were not included at all. Those are all fairly text heavy by any reasonable standard, so I assume they were "punished" because they use JS? Would make sense that nearly all of them are way out of date.

Then I searched specifically for "Baldur's Gate Wiki" but still out of luck - some results, but nothing vaguely Wiki-like.

Finally I searched for "Baldur's Gate Fandom Wiki". This is basically "search engine easy mode", by giving essentially the name of of the site I am looking for. I got ZERO results. At this point I gave up and decided that this thing is useless.

Look, I'm all for unearthing good long-form content (in fact I would say that much of the content around this specific game would qualify), and I do get as annoyed at modern SPAs as the next grumpy neckbeard.

I think considering both of those in a search engine is not a bad idea in and of itself. But I have to wonder what's the point of a search engine that weights some arbitrary aspect of web design higher than the relevancy of the subject matter (to the point of not returning any results at all)? In fact, considering that generally speaking more recent websites tend to include more scripting, you are intentionally skewing the results towards (very) old content, which is probably doing the user a disservice.

Re: A search engine that favors text-heavy sites and punishes modern web design

#474

Pretty cool. I am not sure yet how useful, but cool it is. However, it seems that it currently does not support non-Latin alphabets. Which I understand in an early version. Still, it's handling of such "exception cases" could be improved: when I search for a Russian word, say "Аквариум", I get >, which is rather rude...

"It also focuses on websites in English, Swedish and Latin and tries to identify and ignore the rest (best-effort)." https://news.ycombinator.com/item?id=28551183

Fine; still "Not a supported language" would be much better response than "not a word".

Re: A search engine that favors text-heavy sites and punishes modern web design

#475

Wow, that's awesome. Great work! For a simple test, I searched "fall of the roman empire". In your search engine, I got wikipedia, followed by academic talks, chapters of books, and long-form blogs. All extremely useful resources. When I search on google, I get wikipedia, followed by a listicle "8 Reasons Why Rome Fell", then the imdb page for a movie by the same name, and then two Amazon book links, which are totall…

Search engines whose revenue is based on advertising will ultimately be tuned to steer you to the ad foodchain. All the incentives are aligned towards and all the metrics ultimately in service of, profit for advertisers. Not in the 99% of people who can convinced to consume something by ads? Welp, screw you.

Re: A search engine that favors text-heavy sites and punishes modern web design

#477

Yeah so this is my project. It's very much a work in progress, but occasionally I think it works remarkably well for something I cobbled together alone out of consumer hardware and home-made code :-)

Love the idea. A little feedback: layout needs tweaking for mobile. FWIW: I'm on mobile Firefox for Android.

Re: A search engine that favors text-heavy sites and punishes modern web design

#479

Earlier quoted context omitted.

Good comparison. Reminds me of an analogy I like to make of today's web, which is it feels like browsing through a magazine store — full of top 10s, shallow wow-factoids, and baity material. I genuinely believe terrible results like this are making society dumber.

what I really want is a true AI to search through all that and figure out the useful truth. I don't know how to do this (and of course whoever writes the AI needs to be unbiased...)

>whoever writes the AI needs to be unbiased...)

I'm not sure the idea of a sentient being not having a bias is meaningful. Reality, once you get past the trivial bits, is subjective.

Re: A search engine that favors text-heavy sites and punishes modern web design

#480

Earlier quoted context omitted.

Thanks for the advice; not hacked, but I have "resurrected" many WP sites that have been (including my wife's non-profit). Just running on an EC2 micro instance, but I tried adding "site:" and received "No such domain". Actually, I think it's because I haven't enabled "HTTPS" yet! That's on my to-do along with migrating off EC2-Classic to VPC...

Vanilla HTTP should be fine. I think 80% of the urls are HTTP. If you're getting no such domain, it's either blocked because it looks too much like a spam domain, or it simply hasn't been discovered yet. What's the TLD? I severely restrict some cheaper TLDs because they gave so much spam. For example, cr.yp.to is an example of a baby I know I've definitely thrown out with the bathwater.

Is a good ol' .com with no ads and minimal JS - originally launched in 2011. Thanks again for your insights; I've bookmarked your site and will check back every so often to see if my site's been indexed.
Post reply on HN