Live data from Hacker News

Matt Cutts is looking for scraper sites

twitter.com

61–70 of 120 posts

Re: Matt Cutts is looking for scraper sites

#61

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

Here's hoping they get it right, because the way it stands now, building a good aggregator is very tricky. The "this is why we can't have nice things" attitude is really putting a damper on valuable developments, and it's very close to stifling competition. Why is it anyone's problem that their algo is not smart enough to see potential value in new kinds of services? Even human vs. software curation should be irrelevant here.

Re: Matt Cutts is looking for scraper sites

#62

Bah, what would I possibly need with a scraped definition that 1) Hasn't been chunked into 20 pieces of varying grammatical structure which are automatically matched to corresponding questions 2) Hasn't been subsequently pasted over a slideshow of completely irrelevant stock photos in bold, white font 3) Isn't accompanied by a grid of ~30 vaguely related questions helpfully linked to similar pages and tastefully deco…

If you use Chrome, there is a "Personal Blacklist" extension that does essentially what the manual blacklist used to do.

Re: Matt Cutts is looking for scraper sites

#63

It's a funny quip but it's getting more attention than the important piece of news it highlights. Google is finally doing something about scraping sites doing better in search results than original creators. Good. Many people don't write for money, to put ads on their website, or as part of some "content marketing" campaign. All they want is a little recognition. A boost in positioning on the SERP means we will be ge…

If this is the case, depending on how deep they follow this philosophy, it's important because the people or places that actually gather the content or lead to the content creation have invested and are investing more resources than those who simply scrape the data - and they the opportunity to recover costs, otherwise you damage that positive system's ability to continue on their successful path. That matters if you care how quickly we reach a better and better world for everyone. The source of where content is generated (in form of data or other) is also a strong signal of value, and that seems to be a leading metric of what Google wants to offer - the highest value search results.

Re: Matt Cutts is looking for scraper sites

#64
post #52

Indeed. Google is constantly doing things they punish other people for. A happy DDG user, who still uses !g too often though.

https://duckduckgo.com/?q=scraper+site

I'm not sure which point you're trying to make. Did you look at op's submission?

Re: Matt Cutts is looking for scraper sites

#65
That SERP should show only one result from Wikipedia instead of two. It should be on top, have a blue title link to Wikipedia, and look like an answer to the user's question. That could be done by a general mechanism that lets every site customize their representation on the SERP, or by a special case for Wikipedia.

Re: Matt Cutts is looking for scraper sites

#67
post #19
post #16

Cue Bing, DuckDuckGo and any other search engine (except Google, of course) being Google-killed for "scraping". It's the perfect plan!

Google recently flagged my content-curation startup as "Pure Spam", even though it only takes small snippets from the original sources, is 100% human curated, and always links back to the true original source. Not only are the curated pages blocked, but the entire domain is blocked as "pure spam". People who use Google to find a domain instead of typing the full URL now can't find it anywhere. These assholes are just…

That's brutal.

Honestly you might consider switching to a new domain given that you haven't launched yet. It can take a long time to get out of the Google doghouse.

Re: Matt Cutts is looking for scraper sites

#68

Earlier quoted context omitted.

Yes, they do; it's not abuse when you're given explicit permission. CC BY-SA means you can do whatever you want with it as long as you attribute the source as specified.

"as long as you attribute the source" danielbarla said that they presented the material under a false name; this goes beyond copying and becomes plagiarism, which I can't imagine is an intended result of the CC license.

Is the source 'User X' or 'StackOverflow'? When you reference CC BY-SA code you don't reference the people who, say, checked it into git but rather the whole repo.

Re: Matt Cutts is looking for scraper sites

#69
post #29
post #15

There are lots of places where Google decides to "help" me, but sometimes I just want search results. Other times, I actually like getting the curated content (e.g. search for "delta 3810"). Is there a way to disable this? EDIT: I should also note that I'm one of those who switched over to DuckDuckGo for privacy reasons, so I don't see these results as often now.

The content they do provide is often so bad it is almost embarrassing. Search for "Russia" and you get a completely useless map and a list of random facts. It may be useful to a child researching geography but for me it is just annoying. I want content that is curated by people who actually understand the subject. I would pay for a search engine designed by someone who understands my industry. The Google algorithm on…

I think most people would be a little unsettled and at least occasionally annoyed if when they googled something as broad as "Russia" it gave them all very specific to their industry.

Re: Matt Cutts is looking for scraper sites

#70

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

One man's "scrapper" is another man's "aggregator". How do you think Google would view my site if I wrapped Wikipedia's content, with back link and ran my own ads alongside that content? I would imagine not very positively. Also, is it okay that a bigger entity scrapes my content just because they send me traffic? You might not want to bite the hand that feeds you, but it still doesn't make it right.

The de facto standard robots.txt is pretty likely to be respected by Google, so it's fairly easy to stop their scraping your site. Yes, it is opt-out, bit I'd expect it to be.

It may be quite frustrating for an upstart to be denied access while Google is explicitly allowed, but that's another matter.

Post reply on HN