Live data from Hacker News

Giving up on Google

robsheldon.com

171–180 of 192 posts

Re: Giving up on Google

#171
post #127

Earlier quoted context omitted.

Did you turn safe search off on DDG? When I gave these queries a go, I noticed that DDG stripped "breast" out of the query. The results seemed much more relevant once safe search was disabled. Not that I'd want to go around with safe search disabled all the time...

Yes, this is a safe search problem. Presumably http://duckduckgo.com/?q=breast+shimming&kp=-1 is what you want. Edit: FWIW, I fixed this: http://duckduckgo.com/?q=breast+shimming should now work as well.

Excellent! Safe search didn't occur to me.

Re: Giving up on Google

#172

Earlier quoted context omitted.

Personally, I wish someone I like, like Colin Percival, would start an email service. I like the way he runs Tarsnap.

The problem with starting an e-mail service is the crazy level of administration and sysadmin headaches it induces merely to be able to reliably send e-mail to 90%+ of endpoints.. let alone anything else. The engineering problem is a pretty interesting one, but the headache of ensuring your users can actually get their mail someplace makes it seem dull or insurmountable to solo developers.

I completely agree. Not only that, but users will typically deal with lots of software annoyances, but one thing they simply will not tolerate is late or lost email.

Re: Giving up on Google

#173

Earlier quoted context omitted.

Matt, it's great that you think enough of us here to give direct feedback. I wish I could give you some example searches too, but most of what has bugged me lately relates to stuff which might brush up against NDA terms, so I don't feel comfy giving specifics. Overall I still love Google, I must use the search service 50 times a day so complaints about shortcomings are a bit like grumbling about the paintwork on my f…

Thanks for the feedback. Punctuation is tricky, because it only adds value for power searchers, but indexing it would swell the index size quite a lot. Otherwise, using double-quotes should do an exact search. On content farms, we've definitely heard that feedback. One point up for debate is whether to respond with algorithms-only, or whether we should update our quality guidelines to call out low-quality content far…

Yes!

and if people wanted a fuller result, you can say, "more lower quality links found. Click here to view all" similar to the way that searching in gmail lets you search in trash & spam.

Re: Giving up on Google

#174
My biggest complaint is that queries with something like OpenBSD don't have the word OpenBSD on the page. I have resorted to adding the names of configuration files I know need to be changed in my queries (e.g. pf.conf).

I do really wish I could have a "NO" button when it asks "Did you mean this?". It might give some feedback.

Re: Giving up on Google

#175

As far as I can tell Google is basically giving up on the search business. About a month ago I reported a couple spam sites that were in the top ten results for a fairly popular search phrase. Over a month later they're still there. These are sites that are literally just a list of keywords with no actual content of value on the site, and Google does nothing even though they come up at the top of the results for a ph…

Alex3917, do you remember the sites (or the query) you did?

"learn to hack"

The top result (cyber-trace.com) is just a list of keywords if you scroll down. Also on the next page of results, learn-to-hack.com is the same thing.

For what it's worth, the reason why I submitted this is that I have the page squidoo.com/hacking which seems to alternate between being on the front page of Google's results for that search and being completely missing from the index, depending on what week it is. That might have to do with Squidoo's sitemap and not just Google, but it's certainly peculiar.

Re: Giving up on Google

#176

Earlier quoted context omitted.

Actually I remove tons of spam from the Yahoo, Bing & other feeds.

What order of magnitude is tons? Not to take away from what you're doing, but historically 90% of new domains are spam.

I'd have to check for exact numbers, but for a large % of searches I'm removing links from those APIs.

Re: Giving up on Google

#177

I left a few comments below, but I wanted to say thanks for mentioning some searches that Google didn't do well on. They're interesting searches, so I thought I'd break them down: - [avaya 103r manual] It's a fair complaint to say some low-quality results are returned, but there's a reason. Do that search and Google says "About 1,510 results." That's a minuscule number of results--it usually means the web has very li…

Matt, it's great that you think enough of us here to give direct feedback. I wish I could give you some example searches too, but most of what has bugged me lately relates to stuff which might brush up against NDA terms, so I don't feel comfy giving specifics. Overall I still love Google, I must use the search service 50 times a day so complaints about shortcomings are a bit like grumbling about the paintwork on my f…

I don't want ti speell-checked

/facepalm

Re: Giving up on Google

#178
post #125

Earlier quoted context omitted.

I have given feedback on search for "BMTC" not returning the official website of BMTC (bmtcinfo.com). (2-3 months back). But the search results haven't changed since then.

I did a quick check, and that website has some issues. Here's a wget on bmtcinfo.com: $ wget http://bmtcinfo.com/ --2010-09-14 09:01:43-- http://bmtcinfo.com/ Resolving bmtcinfo.com... failed: Name or service not known. wget: unable to resolve host address `bmtcinfo.com' The url http://bmtcinfo.com/ just doesn't work in a browser or with wget. And trying to fetch the "www" version of the website, it does a 302 redire…

Thanks a lot for the reply. I have passed on this info to bmtcinfo.com through their feedback form.

Re: Giving up on Google

#179

Earlier quoted context omitted.

From your mouth to Marissa Meyer's ears. My ultimate Google fantasy: An account setting called "2008 mode." No instant search. No fancy, annoying endless scroll Google image search. No word clustering/auto-substitution. It would be awesome.

For regular search, their SSL search page https://www.google.com/ is very close to "2008 mode". No luck on the image search part though.

For some reason, this is the page I get whwn I type google.com into the address bar in Safari (Webkit nightly). And I have yet to see Google Instant.

Re: Giving up on Google

#180

Earlier quoted context omitted.

Matt, it's great that you think enough of us here to give direct feedback. I wish I could give you some example searches too, but most of what has bugged me lately relates to stuff which might brush up against NDA terms, so I don't feel comfy giving specifics. Overall I still love Google, I must use the search service 50 times a day so complaints about shortcomings are a bit like grumbling about the paintwork on my f…

Thanks for the feedback. Punctuation is tricky, because it only adds value for power searchers, but indexing it would swell the index size quite a lot. Otherwise, using double-quotes should do an exact search. On content farms, we've definitely heard that feedback. One point up for debate is whether to respond with algorithms-only, or whether we should update our quality guidelines to call out low-quality content far…

I know there's no simple answers for these things. On punctuation, consider the (admittedly obscure) situation of searching for some command line string: obscure_utility "-unknown" "-switches" - either the - sign excludes stuff you want, or gets stripped inside the quotes.

Farming-wise, I think you should probably keep all those results in your index, even the ones that are composed of nothing more than your top searches separated by random phrases! even if it's not there now, in future it'll be possible to score page content on whether or not it has semantic value and draw inferences about sites or entire domains that are filled with junk. That will be interesting and useful from security, economic, and scientific points of view. In the meantime people will find useful analyses to run against that 'bad' data in your results which would not be practical if you purged too aggressively.

What I had envisioned (which might be a tad ambitious, but bear in mind that you already have 5% of my local CPU for the asking with the desktop tool installed) is per-user search filtering. I may like sculpture but hate politics, so I would always search for 'statue -liberty', you are the other way around so your searches tend more towards 'liberty -statue'; I would very much like to be able to have complex filters on the client side, either locally or on the client-facing parts of your servers - and not just for spam sites.

DDG takes the approach of allowing regex, which is a neat thing for the people who know enough to want it, and it would be interesting if search patterns and/or selective exclusions (as described above) could be stored locally, as either weighting tables or some sort of white/blacklist - always include wikipedia, never include about.com. I'm already running chrome and using a Nexus one, so perhaps some hashing could take place on my computers rather than increasing the load on your servers.

The other reason besides spam is that lately I find myself wanting to do specialized searches, but I don't know how to specify the bounds of the search space. For example, I'm interested in law. But a lot of legal terms are in popular currency, so even if I search for "theft +legal" I may get tons of results for cheap car alarms or something. It would be fantastic if I could get a large set of results by specifying a large number of domain-specific terms - say, "tortfeasor privity precedent appellate" - and then hash and save that result set as a 'search space'. So then I could do more specific searches and know that my results would mostly come from websites devoted to the subject, with few that mention it only incidentally.

In actual fact, the legal resources searchable via Scholar and Books are fantastic. I just picked law as an example of where you might want to temporarily limit your search set because most people can appreciate the difference between writing specifically about legal topics vs things that just mention the subject in passing. If you want to learn how to write a good disclaimer, "legal disclaimer" is not a good start because every 3rd landing page on the net includes that phrase as boilerplate. If users could save and reuse result sets, we'd get more actionable results, you'd (maybe) get lower server loads but more importantly, every successful hit (where the user doesn't search again or try another result for several minutes) is an implicit vote for the relevance of the result to the set, and thus a valid input to a semantic classifier.

Post reply on HN