Live data from Hacker News

How is search so bad? A case study

svilentodorov.xyz

161–170 of 416 posts

Re: How is search so bad? A case study

#161
post #145

> At any rate, I got annoyed at this point (mentioning for those who couldn’t tell), so I switched to DuckDuckGo. For those who might be misled like I used to be DuckDuckGo is just a proxy for Bing.

This line from the wikipedia article about DDG somewhat contradicts you, but it is rather vague: > DuckDuckGo's results are a compilation of "over 400" sources,[15] including Yahoo! Search BOSS, Wolfram Alpha, Bing, Yandex, its own Web crawler (the DuckDuckBot) and others. Could you elaborate on your comment, I am sincerely interested in learning the details.

You can open up two web pages side by side, search the same term in both Bing and DuckDuckGo, and see many of the results are the same in similar order (at least they were the last time I did this, maybe two months ago). DuckDuckGo appears to get most of its results from Bing.

Re: How is search so bad? A case study

#162
SEO is the new spam. We solved spam pretty well, but it was a very different solution space to what is available for the web:

- Spammers had basically two ways to verify their efficacy – they could either sign up to every provider under the stars and test each email with each of them individually, or they could use the absence of a signal as "proof" of being caught by a filter. But neither of these are very efficient. An SEO expert can simply wait for the search engine to detect their changes and verify the result with two or three search engines quickly and automatically.

- For practical purposes whether an email is spam is answered in a binary form: either it ends up in your spam box or it does not. Removing spam-looking things from search results entirely would be devastating for any site victim of a false positive. And how do you implement the equivalent of a spam box in a search engine in a useable way?

- Spam filtering was implemented in different ways on every mail provider, so the bar to entry was "randomized" and spammers would have to be quite careful to pass the filters on a large subset of providers. ISPs and users currently have nowhere near the resources to implement their own ranking rules, but maybe this could be a solution in the mid to long term with massively cheaper hardware.

Re: How is search so bad? A case study

#163
post #145

> At any rate, I got annoyed at this point (mentioning for those who couldn’t tell), so I switched to DuckDuckGo. For those who might be misled like I used to be DuckDuckGo is just a proxy for Bing.

For those who might be misled like I used to be DuckDuckGo is just a proxy for Bing.

Every time the topic of search comes up on HN, someone always jumps in and says this.

Then there are a bunch of other people who jump in and say that Duck is much more than that.

So, which is correct?

Re: How is search so bad? A case study

#164
post #103

Earlier quoted context omitted.

It giving the YC startup running the search backend some visibility doesn't mean it somehow "doesn't even have seaarch".

Exactly, there's a search box on the web site. How it's implemented is an implementation detail. Given that it happens to be something by a company (Algolia) selling this as a SAAS solution, I don't think this is a great advertisement for them either.

Algolia is a YC company.

Re: How is search so bad? A case study

#165

All search is being devoured by SEO

This. For ever engineer that somehow works with search quality, there are thousands of experts who are working to subvert SERPs in some fashion. Pretty sure that if we gained true knowledge about the challenges that search faces due to abuse, it would be like facing one of Lovecraft's cosmic horrors.

Re: How is search so bad? A case study

#166

Earlier quoted context omitted.

The Common Crawl is a thing already. Unfortunately, a "full" text crawl of the internets is a YUUUGE amount of data to manage, and I can't think of anything that could change that in the foreseeable future. That's why I think providing a federated Web directory standard, ala ODP/DMOZ except not limited to a single source, would be a far more impactful development.

Unfortunately, a "full" text crawl of the internets is a YUUUGE amount of data to manage Maybe instead of a problem, there is an opportunity here. Back before Google ate the intarwebs, there used to be niche search engines. Perhaps that is an idea whose time has come again. For example, if I want information from a government source, I use a search engine that specializes in crawling only government web sites. If I w…

That's the problem that web directories solve. It's not that you're wrong, it's just largely orthogonal to the problem that you'd need a large crawl of the internets for, i.e. spotting sites about X niche that you wouldn't find even from other directly-related sites, and that are too obscure, new, etc. to be linked in any web directory.

Re: How is search so bad? A case study

#167

I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. Results are not sorted based on relevance or quality but by SEO experts' efforts at making the search results favor themselves. I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surf…

“I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surface-level knowledge that I get from competing websites who just want to make money off pageviews.”

Is it possible that there is no site providing non fluffy content on your query? For a lot of niche subjects, there really are very few if any substantial content on that topic.

Re: How is search so bad? A case study

#168

I feel like Google has often turned strict commands into fuzzy searching, maybe for a decade? I never heard a clear explanation as to why, I just imagined that it was some sort of A/B tested paternalism. Maybe most users really want fuzzy searches when using the commands I use for a strict search.

I think it's simply a human bias in action - people don't realize when their queries benefit from the "fuzzy matching", and they only notice/remember when they don't get what they want from search and then (often mistakenly) blame fuzzy matching for it as that's what's visible to them.

Re: How is search so bad? A case study

#169
post #55

Earlier quoted context omitted.

But those pretend complaints aren't his complaints. His complaint is "why does this archived reddit page from six years ago without any updates come up on search results for 'things within the past month'?" Which is... reasonable.

It is reasonable. It is also likely that whatever meta information reddit is sending back (in headers or tags) is probably not dated correctly for the time of the origin post. Google COULD offer more time machine features and perform diffing on pages. But a reddit "page" will always have content changes, as everything is generated from a database and kept fresh on the page. The ONLY metric therefore Google could use…

It is also likely that whatever meta information reddit is sending back (in headers or tags) is probably not dated correctly for the time of the origin post.

That doesn't explain why Google lists the old search results as being from this month, while Duck correctly lists them as being from years past.

Re: How is search so bad? A case study

#170
post #51

Earlier quoted context omitted.

> A new breakthrough heuristic today will look something totally different, just as meritocratic and possibly resistant to gaming. I wonder how much of this could be obtained back by penalizing: 1. The number of javascript dependencies 2. The number of ads on the page, or the depth of the ad network This might start a virtuous circle, but in the end, this is just a game of cat-and-mouse, and website might optimize fo…

Given that Google makes money off the ads, that would be hard. DuckDuckGo could pull it off. You need another revenue stream though.

Google has taken on so many markets that I don't think they can do anything reasonably well (or disruptive) without conflicting interests. A breakup is overdue: if they didn't control both search and ads, the web would be a lot better nowadays. If they didn't control web browsers as well, standards would be much more important.
Post reply on HN