Live data from Hacker News

A search engine that favors text-heavy sites and punishes modern web design

search.marginalia.nu

361–370 of 735 posts

Re: A search engine that favors text-heavy sites and punishes modern web design

#361

Earlier quoted context omitted.

This sort of optimization is why simple recipes are typically found at the end of a rambling pointless blog post now. Still, the best way to break SEO is to have actual competition in the search space. As long as SEO remains focused on Google there is an opportunity for these companies to thrive by evading SEO braindamage.

That sort of recipe blog hasn't happened just for SEO. It's also a bit of a "two audiences" problem: if you are coming to that food blogger from a search you certainly would prefer the recipe first and then maybe any commentary on it below if the recipe looks good. If you are a regular reader of that food blogger you are probably invested in the stories up top and that parasocial connection and the recipes themselves…

> If you are a regular reader of that food blogger

I think this assumes facts not in evidence. It certainly seems like an overwhelming number of "blogs" are not actual blogs but SEO content farms. There's no regular readers of such things because there's no actual authors, just someone that took a job on Fivver to spew out some SEO garbage. Old content gets reposted almost verbatim because new results better according to Google.

The only reason these "blogs" exist is to show ads and hopefully get someone's e-mail (and implied consent) for a marke....newsletter.

Re: A search engine that favors text-heavy sites and punishes modern web design

#362

Earlier quoted context omitted.

As long as few people use it, it will be great. Rest assured that the moment it becomes popular, the people who want to game it will appear.

I don't think the existing media-heavy websites are gaming Google to rank higher. It's that Google itself prefers media heavy content; they don't have to "game" anything. I also think a search engine like this would be quite hard to game. An ML-based classifier trained on thousands of text-heavy and media-heavy screenshots should be quite robust and I think would be very hard to evade, so the "game" will become more…

With the advances in text generation by machines that looks, but isn't quite accurate (aka GPT-3), seems like it would be easily gamed (given access to GPT-3). Even without GPT-3, if the content being prioritized is mere text, I'm sure that for a pile of money, I could generate something that looks like Wikipedia, in the sense that it's a giant pile of mostly text, but it would make zero sense to a human reader. (Building an SEO farm to boost ranking of not-wikpedia is left as an exercise for the reader.)

Re: A search engine that favors text-heavy sites and punishes modern web design

#363

Earlier quoted context omitted.

Pretty neat!!! You may already be aware of this, but the page doesn't seem to be formatted correctly on mobile. The content shows in a single thin column in the middle.

Hmm, which OS? I only have a single Android phone so I've only fixed the CSS for that.

I was seeing it on Android w/ Firefox. Seems like it's fixed now though. :)

Re: A search engine that favors text-heavy sites and punishes modern web design

#364

Earlier quoted context omitted.

Definitely; We could create a meta search engine that queries them all, in desktop application format. Let's name it after a famous old scientist, and maybe add the year to prove it's modern: Galileo 2021.

Meta search engines leave a bad taste in everyone's mouth because they've always failed. Here is why https://en.wikipedia.org/wiki/Arrow%27s_impossibility_theore... You can't combine a few different ranked lists and expect to get results better than any of the original ranked lists.

That's an invalid application of this theorem. (It doesn't necessarily hold)

Suppose there's an unambiguous ranked preference by all people among a set (webpages, ranking). Suppose one search engine ranks correctly the top 5 results and incorrectly the next 5 results, while another ranks incorrectly the top 5 and correctly the next 5.

What can happen is that some there may be no universally preferred search engine (likely). In practice, as another commenter noted, you can also have most users prefer more a certain combination of results (that's not difficult to imagine, for example by combining top independent results from different engines for example).

Re: A search engine that favors text-heavy sites and punishes modern web design

#365

Absolute textgasm. Wonder how 'text-only first' prioritization is being implemented, algorithmically speaking?

This post - https://news.ycombinator.com/item?id=28551183 - suggests it's a simple set of hueristics, looking for things like javascript, link/SEO spam, language, amount of text content, etc, filtering out unwanted results and only indexing wanted ones.

Re: A search engine that favors text-heavy sites and punishes modern web design

#366

Earlier quoted context omitted.

As long as few people use it, it will be great. Rest assured that the moment it becomes popular, the people who want to game it will appear.

If there were a wider variety of popular search engines, with different ranking criteria, would sites begin to move away from gaming the system? Surely it would be too hard to game more than one search engine at a time?

It would be a matter of numbers anyway about which they optimize for. A/B testing is already in place and doesn't care about where it comes from, just which one does better.

Re: A search engine that favors text-heavy sites and punishes modern web design

#367

"Don't be afraid to scroll down in the search results, unlike in many other search engines, depending on what you are looking for, you may find the best results in the middle of the listing." This is a very polite way of saying "this engine isn't very good" Overall impressed with the project but I thought the word play there was funny

I felt I needed to add it to help people taught by other search engines that they only get 1-2 good results, and the rest is useless. The reason I'm providing a hundred results is that there are often a lot of results to choose from. If the point is to find something unexpected, and that indeed is the entire point, then that is the only sane design choice.

Like you search for something on Google and similar, and you know what you are going to find. They are so good at searching the Internet and predicting what you are going to click on that you never see something new.

It's a great feat of engineering, but a huge tragedy, because discovering new things, outside of what you our your demographic has previously demonstrated an interest in, it can be absolutely life changing.

Re: A search engine that favors text-heavy sites and punishes modern web design

#368

Wow, that's awesome. Great work! For a simple test, I searched "fall of the roman empire". In your search engine, I got wikipedia, followed by academic talks, chapters of books, and long-form blogs. All extremely useful resources. When I search on google, I get wikipedia, followed by a listicle "8 Reasons Why Rome Fell", then the imdb page for a movie by the same name, and then two Amazon book links, which are totall…

The Wikipedia link at the top is always given. It would maybe be good to make it a little clearer that it's not one of the true results.

Re: A search engine that favors text-heavy sites and punishes modern web design

#369

You should monetise this with amazon affiliate links that are relevant to each search. And then use that money to keep this project going. Google is fantastic, but it has become something different from what it was, the company and the product. It is so refreshing to see a modern tool that encourages exploration of the actual world wide web.

That would be an absolutely awful decision.

Re: A search engine that favors text-heavy sites and punishes modern web design

#370

Earlier quoted context omitted.

> You can't combine a few different ranked lists and expect to get results better than any of the original ranked lists. I am skeptical of this application of the theorem. Here is my proposal: Take the top 10 Google and Bing results. If the top result from Bing is in the top 10 from Google, display Google results. If the top result from Bing is not in the top 10 from Google, place it at the 10th position. You'd have…

Right. Arrow's theorem just says it's impossible to do it in all cases. It's still quite possible to get an improvement in a large proportion of cases, as you're proposing.

I've had jobs tuning up the relevance of search engines with methods like

https://ccc.inaoep.mx/~villasen/bib/AN%20OVERVIEW%20OF%20EVA...

and the first conclusion is "something that you think will improve relevance probably won't"; the TREC conference went for about five years before making the first real discovery

https://en.wikipedia.org/wiki/Okapi_BM25

It's true that Arrow's Theorem doesn't strictly apply, but thinking about it makes it clear that the aggregation problem is ill-defined and tricky. (e.g. note also that a ranking function for full text search might have a range of 0-1 but is not a meaningful number, like a probability estimate that a document is relevant, but it just means that a result with a higher score is likely to be more relevant than one with a lower score.)

Another way to think about it is that for any given feature architecture (say "bag of words") there is an (unknown) ideal ranking function.

You might think that a real ranking function is the ideal ranking function plus an error and that averaging several ranking functions would keep the contribution of the ideal ranking function and the errors would average out, but actually the errors are correlated.

In the case of BM25 for instance, it turns out you have to carefully tune between the biases of "long documents get more hits because they have more words in them" and "short documents rank higher because the document vectors are spiky like the query the vectors". Until BM25 there wasn't a function that could be tuned up properly and just averaging several bad functions doesn't solve the real problem.

Post reply on HN