Earlier quoted context omitted.
Imagine if you were looking for the movie.
I tend to prefer Wikipedia for movies. The exception is actor headshots if I'm trying to identify someone, which Wikipedia lacks for licensing reasons, but otherwise Wikipedia tends to be better than IMDB for most needs. Wikipedia has an IMDB link on every article anyway. Another need I guess might be reviews, for which RT or MC are better than IMDB: not sure if either of those two will fare better than IMDB in this…
A search engine that favors text-heavy sites and punishes modern web design
471–480 of 735 posts
Re: A search engine that favors text-heavy sites and punishes modern web design
#472The results were all either barely relevant, outdated (sites that covered the game back in the 90s/2000s before it was re-released), at best tangentially relevant or complete garbage noise. Some of the most highly relevant pages (such as the Steam store listing, the fandom wiki, the publisher/developer's forums for the re-releases, the Baldur's Gate 3 website and the subreddit) were not included at all. Those are all fairly text heavy by any reasonable standard, so I assume they were "punished" because they use JS? Would make sense that nearly all of them are way out of date.
Then I searched specifically for "Baldur's Gate Wiki" but still out of luck - some results, but nothing vaguely Wiki-like.
Finally I searched for "Baldur's Gate Fandom Wiki". This is basically "search engine easy mode", by giving essentially the name of of the site I am looking for. I got ZERO results. At this point I gave up and decided that this thing is useless.
Look, I'm all for unearthing good long-form content (in fact I would say that much of the content around this specific game would qualify), and I do get as annoyed at modern SPAs as the next grumpy neckbeard.
I think considering both of those in a search engine is not a bad idea in and of itself. But I have to wonder what's the point of a search engine that weights some arbitrary aspect of web design higher than the relevancy of the subject matter (to the point of not returning any results at all)? In fact, considering that generally speaking more recent websites tend to include more scripting, you are intentionally skewing the results towards (very) old content, which is probably doing the user a disservice.
Re: A search engine that favors text-heavy sites and punishes modern web design
#473Re: A search engine that favors text-heavy sites and punishes modern web design
#474Pretty cool. I am not sure yet how useful, but cool it is. However, it seems that it currently does not support non-Latin alphabets. Which I understand in an early version. Still, it's handling of such "exception cases" could be improved: when I search for a Russian word, say "Аквариум", I get >, which is rather rude...
"It also focuses on websites in English, Swedish and Latin and tries to identify and ignore the rest (best-effort)." https://news.ycombinator.com/item?id=28551183
Re: A search engine that favors text-heavy sites and punishes modern web design
#475Wow, that's awesome. Great work! For a simple test, I searched "fall of the roman empire". In your search engine, I got wikipedia, followed by academic talks, chapters of books, and long-form blogs. All extremely useful resources. When I search on google, I get wikipedia, followed by a listicle "8 Reasons Why Rome Fell", then the imdb page for a movie by the same name, and then two Amazon book links, which are totall…
Re: A search engine that favors text-heavy sites and punishes modern web design
#476Re: A search engine that favors text-heavy sites and punishes modern web design
#477Yeah so this is my project. It's very much a work in progress, but occasionally I think it works remarkably well for something I cobbled together alone out of consumer hardware and home-made code :-)
Re: A search engine that favors text-heavy sites and punishes modern web design
#478Re: A search engine that favors text-heavy sites and punishes modern web design
#479Earlier quoted context omitted.
Good comparison. Reminds me of an analogy I like to make of today's web, which is it feels like browsing through a magazine store — full of top 10s, shallow wow-factoids, and baity material. I genuinely believe terrible results like this are making society dumber.
what I really want is a true AI to search through all that and figure out the useful truth. I don't know how to do this (and of course whoever writes the AI needs to be unbiased...)
I'm not sure the idea of a sentient being not having a bias is meaningful. Reality, once you get past the trivial bits, is subjective.
Re: A search engine that favors text-heavy sites and punishes modern web design
#480Earlier quoted context omitted.
Thanks for the advice; not hacked, but I have "resurrected" many WP sites that have been (including my wife's non-profit). Just running on an EC2 micro instance, but I tried adding "site:" and received "No such domain". Actually, I think it's because I haven't enabled "HTTPS" yet! That's on my to-do along with migrating off EC2-Classic to VPC...
Vanilla HTTP should be fine. I think 80% of the urls are HTTP. If you're getting no such domain, it's either blocked because it looks too much like a spam domain, or it simply hasn't been discovered yet. What's the TLD? I severely restrict some cheaper TLDs because they gave so much spam. For example, cr.yp.to is an example of a baby I know I've definitely thrown out with the bathwater.