Live data from Hacker News

Google search and search engine spam

googleblog.blogspot.com

131–140 of 223 posts

Re: Google search and search engine spam

#131
post #72

Earlier quoted context omitted.

I don't think they're talking about duplicate content so much as stuff like Demand Media (ehow), Associated Content, Hubspot, ezinearticles, etc, etc.

What's wrong with eHow?

This is HN, so I'll give an example related to my startup. We make a wireless (802.11g) flash drive called the AirStash, which works like an ordinary USB SD card reader on a PC and uses HTML5 for the interface on wireless devices. Here's an example of spam content from eHow that talks about our product:

http://www.ehow.com/how_6861903_install-wireless-flash-drive...

"Wireless flash drives communicate with wireless devices using wireless protocols."

You don't say?

"Advanced wireless flash drives stream data to more than one device at a time."

Well, we're the only one out there, so I guess they're all advanced!

"Insert the USB portion of your wireless flash drive into an available USB port on your computer. Your computer should automatically recognize the device. If not, click the "Start" button and then click "My Computer." Double click the flash drive in the removable media section. This opens the drive and displays the files. Double click the executable file (start.exe, for example) to start the installation process."

Uh, no. No software installation is ever required, a point we make abundantly clear on our web site.

Basically, it's all wrong.

Re: Google search and search engine spam

#132

Earlier quoted context omitted.

Agreed. My ideal search engine wouldn't require real websites to play in the SEO arms-race to beat out the junk sites.

That ideal search engine would find itself quickly the target of people that would try to gain an advantage by figuring out how it works. And then another SEO cycle would start. Don't forget that before google came along nobody was trying to 'game the system' with backlinks and other trickery, the fact that that google is successful is what caused people to start gaming google.

If it were "ideal", it wouldn't be game-able. I'm not going to claim that this ideal is possible!

Re: Google search and search engine spam

#133
post #107

Earlier quoted context omitted.

Because that wouldn't solve the problem for clones of other sites, or clones in other languages. And the Stack Overflow cloners could just make other websites. That's why a primary instinct in search quality is to look for an algorithmic solution that goes to the root of the problem. That approach works across different languages, sites, and if someone makes new sites. To be clear: the webspam team does reserve the r…

> Because that wouldn't solve the problem for clones of other sites, or clones in other languages. Sites that are the victims of content cloning have to be very visible and valuable, so maybe a little manual curating could be relevant. > the Stack Overflow cloners could just make other websites Not really? The point is not to tag the clones but to tag the original; everything that is not the original and that has cop…

this was discussed in an earlier thread and it seemed like the idea of finding the "original" gets really messy and game-able (and potentially oppressive if curation were used).

the primary input to search engines comes from web crawlers...the idea of "first" when it comes to duplicated content is already difficult to determine, and (I would guess) it would get much much worse in the inevitable arms race if something like this were implemented.

Re: Google search and search engine spam

#134
post #49

Earlier quoted context omitted.

Google knows what that site is. If they're serious about their standards, they would remove Mahalo en masse from their index. edit: Or, to satisfy lukev, they can keep the index as-is but make sure Mahalo pages never rank high in results.

It might be better to cripple the site rather than kill it. If all of Mahalo's pages disappeared you can be sure they would return en masse when they found the workaround for the filter. Blocking chunks of their content might make finding the workaround harder and may ultimately force them (or any other low quality site) to up the quality - yeah, I know I am deluding myself.

Exactly; it's not like domain names or IPs are expensive.

Re: Google search and search engine spam

#135
post #110

Earlier quoted context omitted.

I guess if you find it helpful, nothing. I always skip their results because they're always really shallow. The kind of article you'd expect if you were paying someone $4 to research and write an article on a topic they know nothing about.

They are content farms - plain and simple. You might consider throwing about.com in there, too. They're made to earn the employers revenue, not help out people with the particular queries.

The about.com example shows how hard distinguishing quality algorithmically gets at the margins. You do actually find stuff on about.com now and then which is pretty decent--not great but decent. You might even say the same of ehow, albeit at a much lower rate.

So you don't want to just say "these are spam."

Re: Google search and search engine spam

#136
post #110

Earlier quoted context omitted.

I guess if you find it helpful, nothing. I always skip their results because they're always really shallow. The kind of article you'd expect if you were paying someone $4 to research and write an article on a topic they know nothing about.

They are content farms - plain and simple. You might consider throwing about.com in there, too. They're made to earn the employers revenue, not help out people with the particular queries.

The about.com example shows how hard distinguishing quality algorithmically gets at the margins. You do actually find stuff on about.com now and then which is pretty decent--not great but decent. You might even say the same of ehow, albeit at a much lower rate.

So you don't want to just say "these are spam."

Re: Google search and search engine spam

#137

Earlier quoted context omitted.

That ideal search engine would find itself quickly the target of people that would try to gain an advantage by figuring out how it works. And then another SEO cycle would start. Don't forget that before google came along nobody was trying to 'game the system' with backlinks and other trickery, the fact that that google is successful is what caused people to start gaming google.

If it were "ideal", it wouldn't be game-able. I'm not going to claim that this ideal is possible!

Any real-world search engine is going to be analyzed until enough of its internal mechanisms are laid bare to allow gaming to some extent.

Typically you pretend the search engine is a black box, you observe what goes in to it (web pages, links between them and queries) and you try to infer its internal operations based on what comes out (results ranked in order of what the engine considers to be important).

Careful analysis will then reveal to a greater or lesser extent which elements matter most to it and then the gaming will commence. Only by drastically changing the algorithm faster than the gamers can reverse-engineer the inner workings would a search engine be able to keep ahead but there are only so many ways in which you can realistically speaking build a search engine with present technology.

Your ideal, I'm afraid, is not going to be built any time soon, if you have any ideas on how to go about this then I'm all ears.

Re: Google search and search engine spam

#138
post #88

Earlier quoted context omitted.

I've been tracking how often this happens over the last month. It's gotten much, much better, and one additional algorithmic change coming soon should help even more. I'm not saying that a clone will never be listed above SO, but it definitely happens less often compared to a several weeks ago.

While searching for a pdf file, I hardly find the pdf through the top results. The top results are often kind off pdf, ebook search engines themselves and they clearly appear on top by gaming Google. I hope this gets fixed too.

If you know it's a pdf that you're looking for, you can add filetype:pdf to your query.

Re: Google search and search engine spam

#139
post #47

One misconception that we’ve seen in the last few weeks is the idea that Google doesn’t take as strong action on spammy content in our index if those sites are serving Google ads. That's not quite what I've been reading. I believe the more common claim is that Google has a disincentive to algorithmically weed out the kind of drivel that exists for no other reason than to make its publisher money via AdSense. It's abo…

I agree. Also interesting to see that Google defines webspam as "pages that cheat" or "violate search engine quality guidelines." By this definition, scraper sites are not spam at all. Nor are the spammy sites in my field which super-optimize for keywords in ways that make it difficult for legitimate content to rise to visibility. If Google did not operate AdSense, it seems hard to believe the company would not have…

"By this definition, scraper sites are not spam at all."

Disagree. Our quality guidelines at http://www.google.com/support/webmasters/bin/answer.py?hl=en... say "Don't create multiple pages, subdomains, or domains with substantially duplicate content." Duplicate content can be content copied within the site itself or copied from other sites.

Stack Overflow is a bit of a weird case, by the way, because their content license allowed anyone to copy their content. If they didn't have that license, we could consider the clones of SO to be scraper sites that clearly violate our guidelines.

Re: Google search and search engine spam

#140
post #68

Earlier quoted context omitted.

Google has taken action on Mahalo before and has removed plenty of pages from Mahalo that violated our guidelines in the past. Just because we tend not to discuss specific companies doesn't mean that we've given them any sort of free pass.

On a similar note, how is the expert sex change site still in your index? They very clearly are serving different content to the crawler (as evidenced by the "cached" link) than they are to people who click through on the SERPs. I though this was a big no-no? For an example (which was submitted as search feedback a month ago), try searching for "XMPP load balancing" and look at the third organic link. (Edit: actually…

As your edit indicates, when we've looking into Experts Exchange, they weren't cloaking--they were showing the same content to users that they show to Googlebot. If they were cloaking, that would be a violation of our guidelines and thus grounds for removal from our search results.
Post reply on HN