Live data from Hacker News

Matt Cutts is looking for scraper sites

twitter.com

101–110 of 120 posts

Re: Matt Cutts is looking for scraper sites

#101
post #53

Google is taking Cognitive Dissonance to a new level: It's okay for them to scrape every single site, download its content and images, and cache it on their servers, and run their AD platform on top of it. but that's not enough, they would still like to impose their rules and punish people who do the same thing.

You can tell google not to download your site with your robots.txt file.

Also google can impose whatever rules they fancy, because it's their own site. Do you have a website? If I find some that you govern it with some rule that I don't like, should I rage on forums about it? Should you care if I rage about it?

Re: Matt Cutts is looking for scraper sites

#102
post #96

Google is most certainly crossing the line here. 1. They are not only doing this with wikipedia, but with many, many sites: "what is the smallest cell in the human body", "what is the biggest planet in the solar system". 2. The sites they chose to link are not always the highest quality sites, such as the two examples above- why are these websites being featured? 3. Many times, the user will get their answer right th…

All of your arguments are based on the underlying assumption that being a "pure web search engine" is inherently better than their current striving towards being an "knowledge engine" modeled after the Star Trek computer (of which web results are a just a subset). I'm not sure that can be taken as a given, if only because the later presents a much clearer model/metaphor in the mobile first technology climate.

Re: Matt Cutts is looking for scraper sites

#103

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

"Google is finally doing something about scraping" I hope this is genuine and not a disingenuous diversion on Google's part. The fact that the Huffington Post still ranks very high for trendy searches makes me wonder. As usual, follow the money: the scraping sites exist to make money, often through Google's advertising; Google gets a cut. The original content is often on sites with no advertising or real traffic, fro…

This is not Google activing on the community's behalf. It's Google doing a little CYA. It's one of those situations where Google's PR is trying to quell problems that dings their bottom line of advertising revenue. This was a mirror of the FB "official" advertising vs. "bot" advertising through things like Fiver.com. Still a sham on both sides rather than really make a big stink about it.

Re: Matt Cutts is looking for scraper sites

#104
post #46
post #35

Earlier quoted context omitted.

Curious - Which site is that?

You won't find it on Google :) We're making a products recommendation site, focusing on goal-oriented decisions. We'll make the full announcement within the next couple of weeks. This reminds me. Google is really killing the "release early and release often" approach, if people will now have to do a ton of SEO learning and tweaking to avoid having your MVP permanently banned at launch day.

Your site sounds really interesting! Look forward to learning more about it.

Re: Matt Cutts is looking for scraper sites

#105
post #53

Google is taking Cognitive Dissonance to a new level: It's okay for them to scrape every single site, download its content and images, and cache it on their servers, and run their AD platform on top of it. but that's not enough, they would still like to impose their rules and punish people who do the same thing.

I'll concede that there's probably a middleground that's kinga gray, but are you really going to defend bona fide scraper sites? Like ones that simply grab all the text off some other site and repost it, adding no value? Google is obviously adding value by providing snippets from Wikipedia.

Re: Matt Cutts is looking for scraper sites

#107

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

Wikipedia's top rankings are actually a big problem. I know of a site that was the first to put up high-quality reference-type content on the Web and for a while getting reasonable traffic from Google. Wikipedia's editors copied that content into thousands of articles in various ways. Thousands with attribution or copying just the facts and thousands without and copying more than just the facts.

This original site is now getting so little traffic from Google that more people visit it from the trickle of these bottom-of-the-page Wikipedia links than from Google itself. Its traffic was also badly hurt by Google's Panda algorithm, which I think clearly proves how flawed it is since this algorithm was supposed to do the exact opposite.

Because of this situation, if somebody thinks of spending money to create high quality reference-type content, I would strongly advise against it. You have no chance vs. Wikipedia's poorly-written articles repurposing your content and Google's flawed algorithms.

Re: Matt Cutts is looking for scraper sites

#108

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

Wikipedia's top rankings are actually a big problem. I know of a site that was the first to put up high-quality reference-type content on the Web and for a while getting reasonable traffic from Google. Wikipedia's editors copied that content into thousands of articles in various ways. Thousands with attribution or copying just the facts and thousands without and copying more than just the facts. This original site is…

It seems a bit odd for you to be so cagey about the identity of this "original site" while at the same time lamenting that they aren't getting the traffic they deserve. Why don't you tell us who they are?

Re: Matt Cutts is looking for scraper sites

#109
post #108

Earlier quoted context omitted.

Wikipedia's top rankings are actually a big problem. I know of a site that was the first to put up high-quality reference-type content on the Web and for a while getting reasonable traffic from Google. Wikipedia's editors copied that content into thousands of articles in various ways. Thousands with attribution or copying just the facts and thousands without and copying more than just the facts. This original site is…

It seems a bit odd for you to be so cagey about the identity of this "original site" while at the same time lamenting that they aren't getting the traffic they deserve. Why don't you tell us who they are?

It's because I don't speak for the owners of the site and I'd rather make sure they don't mind me putting it out there like this. I could let you know privately, if you'd like to check my story for yourself, though.

Re: Matt Cutts is looking for scraper sites

#110
post #46
post #35

Earlier quoted context omitted.

Curious - Which site is that?

You won't find it on Google :) We're making a products recommendation site, focusing on goal-oriented decisions. We'll make the full announcement within the next couple of weeks. This reminds me. Google is really killing the "release early and release often" approach, if people will now have to do a ton of SEO learning and tweaking to avoid having your MVP permanently banned at launch day.

Depending on what products you're talking about, you should see if it competes in any way with Google's paid results on product searches. If there is any crossover, the USDOJ might be happy to hear what you have to say.
Post reply on HN