Live data from Hacker News

Algorithmic search is sinking

skrenta.com

11–20 of 49 posts

Re: Algorithmic search is sinking

#11
post #5

Ironic that he's proposing to have people solve a problem that arose because people were being manipulated, of course your people cannot be manipulated. No, the solution to this problem is that GetSatisfaction et al use rel=nofollow. It's as simple as that. And arguably Google could improve its algorithm by taking negativity into account.

yes, this is all getsatisfaction's fault. surprised nytimes missed that angle.

Their response:

http://blog.getsatisfaction.com/2010/11/28/when-businesses-a...

Re: Algorithmic search is sinking

#12
post #7

Or maybe our algorithms just aren't good enough. Suppose you use bayesian filtering on the text surrounding the links to determine whether the connection is good or bad. With a reasonable amount of data it should be possible. Note: I'm not an algorithms guy, I do business and strategy and a wee bit of programming, so maybe the example isn't good, but I thinkthe point is.

Interesting, do people frequently use bayes' theorem in web programming? Ive only seen it it other programming contexts.

I have no idea...

That's how I'd solve this particular problem though. As I said in the parent I only have cursory experience in programming, and almost none in algorithms.

Re: Algorithmic search is sinking

#14
I don't know if Skrenta's approach is perfect (can spammers make slashtags? I'll bet they can!) but Google's is clearly failing.

Giant swathes of Google searches are now overrun with datafog spammers. Ehow, squidoo, hubpages, wikihow, buzzle, how-wiki, ezinearticles, bukisa, wisegeek, articlesnatch, healthblurbs, associatedcontent - all thee and thousands more domains filled with spam semi-automatically generated by legions of Indians for a few cents per page.

There's not one word of useful information on any of those domains. But apparently they serve a lot of ads for Google, so they don't get delisted.

Re: Algorithmic search is sinking

#15
There is little rigor behind most of the claims of the NYTime story: The targeted site already negates any pagerank benefit of their links (they do implement nofollow), and the definitive example seems to be nothing more than good SEO of the site in question (most of the other front and second page sites are pretty mediocre as well, clearly with little web competition in the keyword space).

In any case, go to a shopping specific (sub)site if shopping. A google search is a terrible way of find either products or retailers.

Re: Algorithmic search is sinking

#17
"The only way to combat this and return trust and quality to search is by taking an editorial stand and having humans identify the best sites for every category."

There are billions of webpages. Who is going to do this review?

Is someone honestly going to review http://stackoverflow.com/questions/4300234/how-might-union-f... and put it in the category of "How Union/Find data structures can be applied to Kruskal's algorithm?"?

No.

The closest thing to a editorialized web is www.dmoz.org, and that hasn't been properly updated in years (and never will be) because it failed.

Search has to be done with algorithms - there are just too many search queries to do it any other way. Udi Manber, Google’s VP of Engineering stated that 20-25% of all queries made each day have never been seen before: http://www.readwriteweb.com/archives/udi_manber_search_is_a_....

Re: Algorithmic search is sinking

#18
post #7

Or maybe our algorithms just aren't good enough. Suppose you use bayesian filtering on the text surrounding the links to determine whether the connection is good or bad. With a reasonable amount of data it should be possible. Note: I'm not an algorithms guy, I do business and strategy and a wee bit of programming, so maybe the example isn't good, but I thinkthe point is.

In this case, what you specifically want is Sentiment Analysis. It's getting pretty accurate+efficient, and should be usable in just this scenario.

Re: Algorithmic search is sinking

#20

I don't know if Skrenta's approach is perfect (can spammers make slashtags? I'll bet they can!) but Google's is clearly failing. Giant swathes of Google searches are now overrun with datafog spammers. Ehow, squidoo, hubpages, wikihow, buzzle, how-wiki, ezinearticles, bukisa, wisegeek, articlesnatch, healthblurbs, associatedcontent - all thee and thousands more domains filled with spam semi-automatically generated by…

I invoke SandGorgon’s law of outsourcing analogies

As an online discussion about PROGRAMMING grows longer, the probability of a comparison involving outsourcing or Indians approaches 1, if Godwin’s law has not already been satisfied

Post reply on HN