Live data from Hacker News

How is search so bad? A case study

svilentodorov.xyz

81–90 of 416 posts

Re: How is search so bad? A case study

#81
post #37

In my opinion, Google is getting worse constantly, which boils down to basically the following aspects for me: 1. I don’t like the UI anymore. I preferred the condensed view, with more information and less whitespace. 2. Popping up some kind of menu when you return from a search results page shifts down the rest of the items resulting in me clicking search links I am not interested in. 3. It tries to be smarter than…

Agreed on all three but especially #1. The "modern" web is full of so much whitespace it's infuriating

Re: How is search so bad? A case study

#82

Google has definitely stopped being able to find the things I need. Pasting stack traces and error messages. Needle in a haystack phrases from an article or book. None of it works anymore. Does this mean they are ripe for disruption or has search gotten harder?

I've been using DDG as a good enough search engine for most things, but when I sometimes fall back to Google, it blows me away how many ads are on the page pretending to be results!

Google images isn't even worth using at all anymore, after that Getty lawsuit that made them remove links to images (the entire damn point of image search as far as I'm concerned..)

Re: How is search so bad? A case study

#83
post #7

Earlier quoted context omitted.

My guess is that suppressing spammy pages got too hard. So they applied some kind of big hammer that has a high false positive rate. You're getting the best of what's left. Maybe also some quality decline in their gradual shift to less hand weighted attributes and more ML.

Two anecdotes: It’s really fascinating. 1. My work got some attention at CES so I tried to find articles about it. Filtering for items that were from the last X days and searching for a product name found pages and pages of plagiarized content from our help center. Loading any one of the pages showed an OS appropriate fake “your system is compromised! Install this update” box. What’s the game here? Is someone trying…

The first seems big right now, on weird subdomains of clearly hacked sites. E.g. some embedded Linux tutorial on a subdomain of a small-town football club.

Re: How is search so bad? A case study

#85

Earlier quoted context omitted.

The backtick in your link broke it: https://en.wikipedia.org/wiki/Goodhart's_law (Were you on mobile and using a smart keyboard?)

I manually added the comma on a mobile smart keyboard. :) Didn't know that doesn't work haha.

That damn ‘Smart Quotes’ misfeature is still causing havoc even after 30 years.

Re: How is search so bad? A case study

#86
I don't agree with the premise of the article, although I accept that the example given is clearly terrible.

I haven't noticed it before, is it a recent bug? It certainly seems a significant one, but not representative of my experience using Google - which I also acknowledge is skewed according to the data they have on you and SEO gaming.

But generally - those constraints acknowledged - I still find Google's search to be one of the modern wonders of the world, still the go to, and - yes - not perfect.

Re: How is search so bad? A case study

#87
post #7

Earlier quoted context omitted.

My guess is that suppressing spammy pages got too hard. So they applied some kind of big hammer that has a high false positive rate. You're getting the best of what's left. Maybe also some quality decline in their gradual shift to less hand weighted attributes and more ML.

My guess is that Google et al are all hell-bent on not telling you that your search returned zero results. They seem to go to great lengths to make sure that your results page has something on it by any means necessary, including: searching for synonyms for words I searched for instead of the specific words I chose, excluding words to increase the number of results (even though the words they exclude are usually the…

> They seem to go to great lengths to make sure that your results page has something on it by any means necessary

You just described how YouTube's search has been working lately. When you type in a somewhat obscure keyword - or any keyword, really - the search results include not only the videos that match, but videos related to your search. And searches related to your keywords. Sometimes it even shows you a part of the "for you" section that belongs to the home page! The search results are so cluttered now.

Re: How is search so bad? A case study

#88

I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. Results are not sorted based on relevance or quality but by SEO experts' efforts at making the search results favor themselves. I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surf…

> However, it will be interesting to figure the heuristics to deliver better quality search results today.

If only there were some kind of analog for effective ways to locate information. Like if everything were written on paper, bound into collections, and then tossed into a large holding room.

I guess it's past the Internet's event horizon now, but crawler-primary searching wasn't the only evolutionary path to search.

Prior to Google (technically: AdWords revenue funding Google) seizing the market, human-currated directories were dominant [1, Virtual Library, 1991] [2, Yahoo Directory, 1994] [3, DMOZ, 1998].

Their weakness was always cost of maintenance (link rot), scaling with exponential web growth, and initial indexing.

Their strength was deep domain expertise.

Google's initial success was fusing crawling (discovery) with PageRank (ranking), where the latter served as an automated "close enough" approximation of human directory building.

Unfortunately, in the decades since we seem to have forgotten how useful hand-currated directories were, in our haste to build more sophisticated algorithms.

Add to that that the very structure of the web has changed. When PageRank first debuted, people were still manually tagging links to their friends' / other useful sites on their own. Does that sound like the link structure we have in the web now?

Small surprise results are getting worse and worse.

IMHO, we'd get a lot of traction out of creating a symbiotic ecosystem whereby crawlers cooperate with human currators, both of whose enriched output is then fed through machine learning algorithms. Aka a move back to supervised web search learning, vs the currently dominant unsupervised.

[1] https://en.m.wikipedia.org/wiki/World_Wide_Web_Virtual_Libra... , http://vlib.org/

[2] https://en.m.wikipedia.org/wiki/Yahoo!_Directory

[3] https://en.m.wikipedia.org/wiki/DMOZ , https://www.dmoz-odp.org/

Re: How is search so bad? A case study

#89
post #7

Google has definitely stopped being able to find the things I need. Pasting stack traces and error messages. Needle in a haystack phrases from an article or book. None of it works anymore. Does this mean they are ripe for disruption or has search gotten harder?

My guess is that suppressing spammy pages got too hard. So they applied some kind of big hammer that has a high false positive rate. You're getting the best of what's left. Maybe also some quality decline in their gradual shift to less hand weighted attributes and more ML.

Well that “big hammer” so to speak is that they tend to favor sites that have a lot of trust and authority.

Someone mentioned that the sites that have the answer typically is buried in the results. That’s because they tend to favor big brands and authoritative sites. And those sites oftentimes don’t have the answer to the search query.

Google’s results have gotten worse and worse over the years.

Re: How is search so bad? A case study

#90
post #23

Earlier quoted context omitted.

I think the Web just kind of stopped being full of searchable information.

Imagine if instead of kneecapping XHTML and the semantic web properties it had baked in, Google had not entered into the web browser space. We might be able to mark articles up with ` `, and set their subject tags to the URN of the people, places, and things involved. We could give things a published and revised date with change logs. Mark up questions, solutions, code and language metadata. All of that is extremely…

Good machine-readable ("semantic") information will only be provided if incentives aren't misaligned against it, as they are on much of the commercial (as opposed to academic, hobbyist, etc.) Web. Given misaligned incentives, these features will be subverted and abused, as we saw back in the 1990s with tags and the like.
Post reply on HN