Live data from Hacker News

A search engine that favors text-heavy sites and punishes modern web design

search.marginalia.nu

611–620 of 735 posts

Re: A search engine that favors text-heavy sites and punishes modern web design

#611
post #607

Wow, that's awesome. Great work! For a simple test, I searched "fall of the roman empire". In your search engine, I got wikipedia, followed by academic talks, chapters of books, and long-form blogs. All extremely useful resources. When I search on google, I get wikipedia, followed by a listicle "8 Reasons Why Rome Fell", then the imdb page for a movie by the same name, and then two Amazon book links, which are totall…

If this search engine ever takes off, the listicle writers will just start optimizing for it too, right?

Mission accomplished, then.

Re: A search engine that favors text-heavy sites and punishes modern web design

#612

Earlier quoted context omitted.

If Google starts showing interesting text-heavy links instead of vapid listicles and storefronts, I have accomplished everything I ever could dream of.

Google Info - for when you're looking for information, not shopping advice or lists!

Maybe you're joking, but this is a good idea for search engine. Better: Credible info.

Re: A search engine that favors text-heavy sites and punishes modern web design

#613

Earlier quoted context omitted.

It’d be nice if you had a page to get the current index status for a domain.

Try a query on the form site:www.example.com ;-)

> site:www.washingtonpost.com

> Blacklisted false

> site:www.wsj.com

> Blacklisted false

> site:www.rt.com

> Blacklisted false

> site:www.nytimes.com

> Blacklisted true

?

Re: A search engine that favors text-heavy sites and punishes modern web design

#614

Earlier quoted context omitted.

Thanks for the advice; not hacked, but I have "resurrected" many WP sites that have been (including my wife's non-profit). Just running on an EC2 micro instance, but I tried adding "site:" and received "No such domain". Actually, I think it's because I haven't enabled "HTTPS" yet! That's on my to-do along with migrating off EC2-Classic to VPC...

Vanilla HTTP should be fine. I think 80% of the urls are HTTP. If you're getting no such domain, it's either blocked because it looks too much like a spam domain, or it simply hasn't been discovered yet. What's the TLD? I severely restrict some cheaper TLDs because they gave so much spam. For example, cr.yp.to is an example of a baby I know I've definitely thrown out with the bathwater.

www.ft.com gets 'no such domain'

Re: A search engine that favors text-heavy sites and punishes modern web design

#615

Earlier quoted context omitted.

But in both cases you face the problem of aggregating preferences of many into one. In one case you are combining personal preferences in the other case aggregating ‘preferences’ expressed by search engines.

But search engines aren't voting to maximize the chances that their preferred candidate shows up on top. The mixed ranker has no requirement to satisfy Arrows integrity constraints. It has to satisfy the end user, which is quite possible in theory. Conditions the mixed ranker doesn't have to satisfy "ranking while also meeting a specified set of criteria: unrestricted domain, non-dictatorship, Pareto efficiency, and…

Sure, but the problem that conventional IR ranking functions are not meaningful other than by ordering leads you to the dismal world of political economy where you can't aggregate people's utility functions. (Thus you can't say anything about inequality, only about Pareto efficiency)

Hypothetically you could treat these functions as meaningful but when you try you find that they aren't very meaningful.

For instance IBM Watson aggregated multiple search sources by converting all the relevance scores to "the probability that this result is relevant".

A conventional search engine will do horribly in that respect, you can fit a logit curve to make a probability estimator and you might get p=0.7 at the most and very rarely get that, in fact, you rarely get p>0.5.

If you are combining search results from search engines that use similar approaches you know those p's are not independent so you can't take a large numbers of p=0.7's and turn that into a higher p.

If you are using search engines that use radically different matching strategies (say they return only p=0.99 results with low recall) the Watson approach works, but you need a big team to develop a long tail of matching strategies.

If you had a good p-estimator for search you could do all sorts of things that normal search engines do poorly, such as "get an email when a p>0.5 document is added to the collection."

For now alerting features are either absent or useless and most people have no idea why.

Re: A search engine that favors text-heavy sites and punishes modern web design

#616
post #303
post #230

Earlier quoted context omitted.

> the websites were all strangely disreputable Interesting you'd feel that way when sites without "modern design" are encountered. Is this your own bias perhaps creating a judgment or are they sites that you already know have a bad reputation?

Or perhaps the websites being returned are garbage? I have the same experience trying a few searches and following the top 5 links. Besides wikipedia, I haven't found a single useful website.

This!

Maybe modern or 'non-modern' web design just isn't a great litmus test for quality content? Could just need some work. At any rate I wasn't clicking on the results.

Re: A search engine that favors text-heavy sites and punishes modern web design

#617
post #322

Earlier quoted context omitted.

The context matters. I'd happily read "Top 10" lists on a website if the site itself was dedicated to that one thing. "Top 10 Prog Rock albums", while a lazy, SEO-bait title, would at least be credible if it were on a music-oriented website. But no, these stories all come from cookie-cutter "new media" blog sites, written by an anonymous content writer who's repackaged Wikipedia/Discogs info into Buzzfeed-style copy…

This got me thinking that maybe one of the other big reasons for this is that the algorithms prioritize newer pages over older pages. This produces the problem where instead of covering a topic and refining it over time, the incentive is to repackage it over and over again. It reminds me of an annoyance I have with the Kindle store. If I wanted to find a book on, let's say, Psychology, there is no option to find all-…

> This got me thinking that maybe one of the other big reasons for this is that the algorithms prioritize newer pages over older pages.

Actually that's not always the case. We publish a lot of blog content and it's really hard to publish new content that replaces old articles. We still see articles from 2017 coming up as more popular than newer, better treatments of the same subject. If somebody knows the SEO magic to get around this I'm all ears.

Re: A search engine that favors text-heavy sites and punishes modern web design

#618
post #280

I tried a few searches. >: none of the search results appeared to have anything to do with Javascript pipe syntax. (Which doesn't exist yet, but it's under discussion.) Google gives a bunch of highly-relevant results. >: first result is a list of books about relativity, one of which is Reichenbach's "Philosophy of space and time"; good, but there's no real information there. Second is about Reichenbach but nothing to…

i don't know, people here like to complain about google, but google still works pretty well for me.

Re: A search engine that favors text-heavy sites and punishes modern web design

#619
post #607

Earlier quoted context omitted.

If this search engine ever takes off, the listicle writers will just start optimizing for it too, right?

Mission accomplished, then.

If the goal was to remove modern web design, ok sure mission accomplished.

If your goal was to create a search engine that ignored listicles and other fluff and instead got you meatier results like "academic talks" and such, then no.

Re: A search engine that favors text-heavy sites and punishes modern web design

#620

Earlier quoted context omitted.

This is exactly how I prefer to use my search engines.

I searched like this all my life and always got expected results. But just a week ago I found out that these "how", "what" questions give better and faster results on Google.

That switch happened some years ago. I've been unlearning and relearning how to use google for what feels like at least three or four years now.

The main pain-point, though, is that a lot of long-tail searches you could've used to find different results in years past, now seem to funnel you to the same set of results based on your apparent intent. At least, it has felt that way -- I'm not entirely sure how the modern google algorithm works.

Post reply on HN