Live data from Hacker News

How is search so bad? A case study

svilentodorov.xyz

121–130 of 416 posts

Re: How is search so bad? A case study

#121

Earlier quoted context omitted.

My guess is that Google et al are all hell-bent on not telling you that your search returned zero results. They seem to go to great lengths to make sure that your results page has something on it by any means necessary, including: searching for synonyms for words I searched for instead of the specific words I chose, excluding words to increase the number of results (even though the words they exclude are usually the…

> They seem to go to great lengths to make sure that your results page has something on it by any means necessary You just described how YouTube's search has been working lately. When you type in a somewhat obscure keyword - or any keyword, really - the search results include not only the videos that match, but videos related to your search. And searches related to your keywords. Sometimes it even shows you a part of…

Searching gibberish to try to get as few results as possible.

I got down to one with "qwerqnalkwea"

"AEWRLKJAFsdalkjas" returns nothing, but youtube helpfully replaces that search with the likewise nonsensical "AEWR LKJAsdf lkj as" which is just full of content.

Re: How is search so bad? A case study

#122

I always wondered why a search engine can't use SEO tactics as a kind of anti-signal. What comes to the top of a google search if you filter out anyone gaming SEO?

Likely because a lot of SEO tactics are not necessarily things that hurt the quality of the site for the user. Using the right heading tags, image alt texts, meta data/schema markup, a good title and meta description etc are all SEO tactics, and they're also all things that help the user experience.

Similarly, a lot of inbound/offsite SEO tactics are theoretically things that help the user as well. Providing content people want to link to, getting authorities to link to said relevant content, etc are all things a user would appreciate.

Using SEO tactics as an anti signal would boost poorly designed sites, inaccessible sites, etc. What really needs to be done is something that filters out sites creating thin content just for the purposes of getting traffic, and that's harder to filter out.

Re: How is search so bad? A case study

#123
post #11

This article is essentially just complaining that DDG and Google don't have special parsing for reddit pages ("How come it doesn't know that thread didn't get many upvotes?", "How come it thinks some change to the site's layout was an update to the page?") Maybe if you want to search reddit, the best search engine is the search bar on reddit.com.

Google is a TRILLION dollar company. Indexing Reddit properly would take what? 2-3 engineers? Cmon.

Re: How is search so bad? A case study

#124

This should probably be a separate submission but why is search so bad everywhere? - Confluence: Native search is horrible IME - Microsoft Help (Applications): .chm files Need I say more. - Microsoft Task Bar: Native search okay and then horrible beyond a few key words and then ... BING :-( - Microsoft File Search: Even with full disk indexing (I turned it on) it still takes 15-20 minutes to find all jpegs with an SS…

> - Microsoft File Search: Even with full disk indexing (I turned it on) it still takes 15-20 minutes to find all jpegs with an SSD. What's going on there?

Does turning it off speed it up? I think disk indexing (the way Windows does it) is a remnant from HDD times, and might make things worse when used together with a modern SSD.

> Adobe PDFs: Readers all versions. What? You mean you want to search for TWO words. Sacrilege. Don't do it.

If you're just viewing and searching PDFs (and don't have to fill out PDF forms on a regular basis), check out SumatraPDF. Fastest PDF reader on Windows I've come accross so far.

Re: How is search so bad? A case study

#125
post #119

I just did a google search for "piano". Just the word "piano" Only one link on the first page, the wikipedia entry for "piano" had anything to do with pianos, (i.e., the instrument invented in Italy 300+ years ago that has hammers, strings, and an iron frame).

What do you get when you search for that? Did a test right now, and apart from the Wikipedia page, I get videos about piano music/pianos, shop pages for buying pianos and local businesses that sell either pianos or piano lessons. So I'm curious whether the issue is that there are too many shopping/business related pages (which is fair, but at least those seem to be piano related), or whether you're getting something…

The first three links after the ad were for the same "virtual piano" (not a piano) on different websites.

See https://imgur.com/a/cMC9wQH

Then the wikipedia page, then a couple of "online" non-pianos, then a company that happens to be called piano.io

https://imgur.com/a/uRcyx84

Shopping pages are fine, if we'd get links to, say Steinway, Yamaha, and Bosendorfer, or links to Lang Lang's home page, or something that has more to do with _pianos_.

Re: How is search so bad? A case study

#126
post #103

Earlier quoted context omitted.

Hacker News doesn't even have search at the moment. It just redirects you to some crappy external site.

It giving the YC startup running the search backend some visibility doesn't mean it somehow "doesn't even have seaarch".

Exactly, there's a search box on the web site. How it's implemented is an implementation detail. Given that it happens to be something by a company (Algolia) selling this as a SAAS solution, I don't think this is a great advertisement for them either.

Re: How is search so bad? A case study

#127
The problem as I see it is that popularity ranking worked fine in the pre Eternal September era for the web (~10 years ago?). I think it is safe to say that most HN users skew toward searching for more technical, intellectual or scientific topics and get frustrated by their searches getting swamped by popular topics. What I'd like to see is a check box or slider bar to exclude or adjust the weighting for popularity in a search. I don't need to see links for the latest Taylor Swift breakup or what the Kardashians are up to that appear in a technical search due to a randomly shared keyword. Often, the topics I am searching for will never be popular and current search operates on the assumption that it will.

A second problem is that now that Google likes to rudely assume to know what you want, i.e. ignoring quotes and negation in search or even modifying keywords, its even harder to find what you want especially if its not on the first page or two of results. Because of this interference even changing your search parameters doesn't change the results much and you see the essentially the same links. What I'd like to see is a search engine that will do a delta between say Google and Bing and drop the links common to the two services. This might lead to uncovering the more esoteric or hidden links buried by the assumptions of the algorithm.

Finally, a last problem that I see right now is the filter bubble effect. I had to search for how to spell "kardashians" in the above paragraph. Now my searches and ads for the next 2-3 weeks will be poisoned by articles or ads about the Kardashians. Taking one for team to make my point, I suppose.

Re: How is search so bad? A case study

#128

Earlier quoted context omitted.

That damn ‘Smart Quotes’ misfeature is still causing havoc even after 30 years.

Nitpick: It's actually an apostrophe " ' ", not a backtick/grave accent " ` " or comma " , " :D https://en.wikipedia.org/wiki/Apostrophe https://en.wikipedia.org/wiki/Grave_accent#Use_in_programmin... https://en.wikipedia.org/wiki/Comma

Oh yes sorry I’m full of mistakes today. Of course, not a comma!

Re: How is search so bad? A case study

#129
post #64
post #51

Earlier quoted context omitted.

> A new breakthrough heuristic today will look something totally different, just as meritocratic and possibly resistant to gaming. I wonder how much of this could be obtained back by penalizing: 1. The number of javascript dependencies 2. The number of ads on the page, or the depth of the ad network This might start a virtuous circle, but in the end, this is just a game of cat-and-mouse, and website might optimize fo…

You could even have all this under one roof: one common search spider that feeds this ensemble of different ranking algorithms to produce a set of indices, and then a search engine front end that round-robins queries out between the different indices. (Don’t like your query? Spin the algorithm wheel! “I’m Feeling Lucky” indeed .)

What if: SEO consultants aren’t gaming system, but the search and web is being optimized for “measurable immediate economic impact” that is ad revenue at this moment — due to web itself being in-monetizable and unable to generate value.

I don’t like the whole concept of SEO, I don’t like the way the web is today, but I think we should stop and think before resorting to “immoral few is destroying things, we unfuck it reclaim what we deserve” type simplification.

Re: How is search so bad? A case study

#130
post #55

Earlier quoted context omitted.

It is reasonable. It is also likely that whatever meta information reddit is sending back (in headers or tags) is probably not dated correctly for the time of the origin post. Google COULD offer more time machine features and perform diffing on pages. But a reddit "page" will always have content changes, as everything is generated from a database and kept fresh on the page. The ONLY metric therefore Google could use…

Seems like an easy solution to this problem would use two functions. One function that takes the output of the page, and renders it so only what's user visible, actually gets indexed. So no headers, no JSON data, no nothing, unless it's actually in the final outcome of the page when rendered. This would require jsdom or some other DOM implementation. Hardly hard for Google (Chrome) to achieve this, and been done mult…

Typically dynamic content doesn't change from second to second, it changes after 5 minutes or an hour or 1 day, actually it is extremely site specific too.

But I do like your idea.

To go a bit further on your idea - you could apply machine learning to analyse the changes. So for example, ML could determine what is probably the "content area" of the page simply by having built out a NN for each website that self-expires the training data at about 1 month (to account for redesigns over time).

The major problem will still be "ads" in the middle of the content, especially odd scroll designs ads that have a different "picture" at each scroll position, as well as video ads that are likely to be different at each screen shot.

Another form of ad being the "linked" words like when you see random words of the paragraphs becoming links that go to shitty websites that define the word but show a bunch of other ads.

I suppose Google could simply install UBlock in it's training data collector harness to help with that stuff. >()

Post reply on HN