Live data from Hacker News

How is search so bad? A case study

svilentodorov.xyz

101–110 of 416 posts

Re: How is search so bad? A case study

#101

I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. Results are not sorted based on relevance or quality but by SEO experts' efforts at making the search results favor themselves. I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surf…

This is very much misguided.

Many websites do have "hacked" (blackhat/shady) SEO, but these websites do not last long, and are entirely wiped out (see: de-ranked) every major algorithm update.

The major players you see on the top rankings today do utilize some blackhat SEO, but it's not at a level that significantly impacts their rankings. Blackhat SEO is inherently dangerous, because Google's algorithm will penalize you at best when it finds out -- and it always does -- and at worst completely unlist your domain from search results, giving it a scarlet letter until it cools off.

However, the bulk of all major websites primary utilize whitehat SEO, i.e "non-hacked," i.e "Google-approved" SEO to maintain their rankings. They have to, else their entire brand and business would collapse, either from being out-ranked or by being blacklisted for shady practices.

Additionally, Google's algorithim hasn't changed much at all from pagerank, in the grand scheme of things. If you can read between their lines, the biggest SEO factor is: how many backlinks from reputable domains do you have pointing at your website? Everything else, including blackhat SEO, are small optimizations for breaking ties. Sort of like PED usage in competitive sports; when you're at the elite level, every little bit extra can make a difference.

Google's algorithm works for its intended purposes, which is to serve pages that will benefit the highest amount of people searching for a specific term. If you are more than 1 SD from the "norm" searching for a specific term, it will likely not return a page that suits you best.

Google's search engine based on virality and pre-approval. "Is this page ranked highly by other highly ranked pages, and does this page serve the most amount of people?" It is not based on accuracy, or informational-integrity -- as many would believe by the latest Medic update -- but simply "does this conform to normal human biases the most?"

If you have a problem with Google's results, then you need to point the finger at yourself or at Google. SEO experts, website operators, etc. are all playing a game that's set on Google's terms. They would not serve such shit content if Google did not: allow it, encourage it, and greatly reward it.

Google will never change the algorithm to suit outliers, the return profile is too poor. So, the next person to point a finger at is you: the user. Let me reiterate, Google's search engine is not designed for you; it is designed for the masses. So there is no logical reason for you to continue using it the way you do.

If you wish to find "deep enough" sources, that task is on you, because it cannot be readily or easily monetized; thus, the task will not be fulfilled for free by any business. So, you must look at where "deep enough" sources lay: books, journals, and experts.

Books are available from libraries, and a large assortment of them are cataloged online for free at Library Genesis. For any topic you can think of, there is likely to be a book that goes into excruciating detail that satisfies your thirst for "deep enough."

Journals, similarly. Library Genesis or any other online publisher, e.g NIH, will do.

Experts are even better. You can pick their brains and get even more leads to go down. Simply, find an author on the subject -- Google makes this very easy -- and contact them.

I'm out of steam, but I really felt the need to debunk this myth that Google is a poor, abused victim, and not an uncaring tyrant that approves of the status quo.

Re: How is search so bad? A case study

#102
This should probably be a separate submission but why is search so bad everywhere?

- Confluence: Native search is horrible IME

- Microsoft Help (Applications): .chm files Need I say more.

- Microsoft Task Bar: Native search okay and then horrible beyond a few key words and then ... BING :-(

- Microsoft File Search: Even with full disk indexing (I turned it on) it still takes 15-20 minutes to find all jpegs with an SSD. What's going on there?

- Adobe PDFs: Readers all versions. What? You mean you want to search for TWO words. Sacrilege. Don't do it.

Seriously though with all the interview code tests bubble sort, quick sort, bloom filters, etc. Why can't companies or even websites get this right?

And I agree with other commenters as far as Google, Bing, DDG, or other search sites it's been going down hill but the speed of uselessness is picking up.

The other nagging problem (at least for me) is that explicit searches which used to yield more relevant results now are front loaded with garbage. If I'm looking for datasheet on an STM (ST Microsystems) Chip and I start search with STM as of today STM is no longer relevant (it is, meaning it shows up after a few pages). But wow it seems like the SEOs are winning but companies that use this technique won't get my business.

Re: How is search so bad? A case study

#103

It's a hard problem because what is relevant is inherently subjective and context specific and only a minority of users uses the advanced search functionality so it is also not a big priority to solve it. Both Google and Duck Duck Go optimize for the simple use case where there's a bit of user context and some short query that the user typed. That's what needs to work well. For that Google is still pretty good. I try…

Hacker News doesn't even have search at the moment. It just redirects you to some crappy external site.

It giving the YC startup running the search backend some visibility doesn't mean it somehow "doesn't even have seaarch".

Re: How is search so bad? A case study

#104
post #51

I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. Results are not sorted based on relevance or quality but by SEO experts' efforts at making the search results favor themselves. I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surf…

> A new breakthrough heuristic today will look something totally different, just as meritocratic and possibly resistant to gaming. I wonder how much of this could be obtained back by penalizing: 1. The number of javascript dependencies 2. The number of ads on the page, or the depth of the ad network This might start a virtuous circle, but in the end, this is just a game of cat-and-mouse, and website might optimize fo…

Given that Google makes money off the ads, that would be hard. DuckDuckGo could pull it off. You need another revenue stream though.

Re: How is search so bad? A case study

#105

I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. Results are not sorted based on relevance or quality but by SEO experts' efforts at making the search results favor themselves. I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surf…

“When a measure becomes a target, it ceases to be a good measure” - Charles Goodhart

Re: How is search so bad? A case study

#106

Earlier quoted context omitted.

I've been using DDG as a good enough search engine for most things, but when I sometimes fall back to Google, it blows me away how many ads are on the page pretending to be results!

Same here, I actually prefer DDG to Google now, even for regional (Germany) results. When I switched, about a year and a half ago, I felt like I was switching to a lesser quality search engine (it was an ethical choice and done because I can), that, however, gradually and constantly got better, whereas Google went the opposite path. Nowdays I only really use Google to leech bandwidth off their maps services. Despite…

> Speaking of bandwidth and OSM reminds me, is there an "SETI-but-for-bandwidth-not-CPU-cycles" kind of thing one could help out with? Like a torrent for map data?

OSM used to have tiles@home, a distributed map rendering stack, but that shut down in 2012. There is currently no OSM torrent distribution system, but I'd like to set that up.

Re: How is search so bad? A case study

#107
I feel like Google has often turned strict commands into fuzzy searching, maybe for a decade?

I never heard a clear explanation as to why, I just imagined that it was some sort of A/B tested paternalism. Maybe most users really want fuzzy searches when using the commands I use for a strict search.

Re: How is search so bad? A case study

#108

I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. Results are not sorted based on relevance or quality but by SEO experts' efforts at making the search results favor themselves. I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surf…

Good to hear your concerns.

> The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left.

I intend to announce the alpha test of my search engine here on HN.

My search engine is immune to all SEO efforts.

> I can possibly not find anything deep enough about any topic by searching on Google anymore.

In simple terms my search engine gives users content with the meaning they want and in particular stands to be very good, by far the best, at delivering content with "deep" meaning.

> I need something better.

Coming up.

> However, it will be interesting to figure the heuristics to deliver better quality search results today.

Uh, sorry, it's not fair to say that my search engine is based on "heuristics".

I'm betting on my search engine being successful and would have no confidence in heuristics.

Instead of heuristics I took some new approaches:

(1) I get some crucial, powerful new data.

(2) I manipulate the data to get the desired results, i.e., the meaning.

(3) The search engine likely has by far the best protections of user privacy. E.g., search results are the same for any two users doing the same query at essentially the same time and, thus, in particular, independent of any user history.

(4) The search engine is fully intended to be safe for work, families, and children.

For those data manipulations, I regarded the challenge as a math problem and took a math approach complete with theorems and proofs.

The theorems and proofs are from some advanced, not widely known, pure math with some original applied math I derived. Basically the manipulations are as specified in math theorems with proofs.

> A new breakthrough heuristic today will look something totally different, just as meritocratic and possibly resistant to gaming.

My search engine is "something totally different".

My search engine is my startup. I'm a sole, solo founder and have done all the work. In particular I designed and wrote the code: It's 100,000 lines of typing using Microsoft's .NET.

The typing was without an IDE (integrated development environment) and, instead, was just into my favorite general purpose text editor KEdit.

It's my first Web site: I got a good start on Microsoft's .NET and ASP.NET (for the Web pages) from

Jim Buyens, Web Database Development, Step by Step, .NET Edition, ISBN 0-7356-1637-X, Microsoft Press.

The code seems to run as intended. The code is not supposed to be just a "minimum viable product" but is intended for first production to peak usage of about 100 users a second; after that I'll have to do some extensions for more capacity. I wrote no prototype code. The code needs no refactoring and has no technical debt.

While users won't be aware of anything mathematical, I regard the effort as a math project. The crucial part is the core math that lets me give the results. I believe that that math will be difficult to duplicate or equal. After the math and the code for the math, the rest has been routine.

Ah, venture capital and YC were not interested in it! So I'm like the story "The Little Red Hen" that found a grain of wheat, couldn't get any help, then alone grew that grain into a successful bakery. But I'm able to fund the work just from my checkbook.

The project does seem to respond to your concerns. I hope you and others like it.

How should I announce the alpha test here at HN?

Re: How is search so bad? A case study

#109
I just did a google search for "piano". Just the word "piano"

Only one link on the first page, the wikipedia entry for "piano" had anything to do with pianos, (i.e., the instrument invented in Italy 300+ years ago that has hammers, strings, and an iron frame).

Re: How is search so bad? A case study

#110

I've been screaming about this for years and only recently have people begun agreeing with me - and I know exactly what the major problems are, and they are synergistic: 1. SEO has totally warped result rankings. Now instead of getting results which naturally match my keywords because of content, I'm presented with almost exclusively commercial websites which are trying to sell me something. Gone are the days where y…

Google has fallen to Goodhart's Law: "When a measure becomes a target, it ceases to be a good measure." For years, Google moved the goalposts for what measurements of a good website were and the broad internet moved in lockstep to meet those goals. Until Google settled on a "thin veneer of content but actually an ad for a service" as not too spammy to blacklist. So that is now the advice for any business. Hey, want to place highly for "St. Louis dentist"? Make a top 10 list of toothbrushing mistakes, etc.

Second, Google shifted about 7-10 years ago from searching for webpages to searching for answers. This was reflected in how they communicated about search and also in stuff like showing infoboxes and AMP results. I think this was move into mobile, but it leaves the actual web underserved.

The worst part is they have been removing tons of older content from the web that is still there and is still valuable but has become no longer surfaceable, even with direct quoted phrases.

Post reply on HN