Live data from Hacker News

Matt Cutts is looking for scraper sites

twitter.com

21–30 of 120 posts

Re: Matt Cutts is looking for scraper sites

#21
Bah, what would I possibly need with a scraped definition that

1) Hasn't been chunked into 20 pieces of varying grammatical structure which are automatically matched to corresponding questions

2) Hasn't been subsequently pasted over a slideshow of completely irrelevant stock photos in bold, white font

3) Isn't accompanied by a grid of ~30 vaguely related questions helpfully linked to similar pages and tastefully decorated with more irrelevant stock photos

4) Only occupies ~1.5 rather than 3 or 4 of the front page search results

5) Contains only closely related textual ads rather than a melange of casino, fast food, and online college banners

6) Has fewer than 25 trustworthy stock faces smiling back at me from any given scroll position

If this is the best google can do then I don't think wiki.answers.com has anything to fear.

------------

Seriously, how the hell does wiki.answers.com manage to pollute half of the searches I make with their algorithmically generated garbage (multiple times, at that)?! What kind of SEO catapulted them to the top despite 0 viewer retention and what surely must be about 0 reputable backlinks? How haven't they been sent to the 1000th page with manual penalties already? They show up before wikipedia itself, for crying out loud!

Google, if you aren't going to let users maintain a manual blacklist, you need to be on top of this kind of thing. It's seriously degrading my search experience and I suspect I'm not alone. This kind of inattention is the type of thing that can push even the most inattentive users to change default search engines.

Re: Matt Cutts is looking for scraper sites

#22
post #12
post #9

Easy. Search for a programming related question. After the result from stackoverflow you'll find dozens of scraper sites.

If you read correctly, the task is to report scraper sites that rank HIGHER than the original site. Not the case in your scenario.

And in OPs context, not only does it rank higher, it's above ALL links. So you don't even need to visit the target page and in affect stealing traffic (and potential ad revenue) from the source.

Re: Matt Cutts is looking for scraper sites

#24
post #9

Easy. Search for a programming related question. After the result from stackoverflow you'll find dozens of scraper sites.

Do we need to differentiate from mechanical scraping versus manual scraping, though? Because realistically a hefty percentage of StackOverflow content are users manually scraping other sites, all in hopes of earning some imaginary internet points.

Re: Matt Cutts is looking for scraper sites

#25

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

Relax, it's a joke.

Re: Matt Cutts is looking for scraper sites

#26
post #20
post #17

Earlier quoted context omitted.

If Google only needs to visit Wikipedia's "scraper" page once a day or less, but serves it out to others with attribution, isn't that helping Wikipedia by lowering traffic COSTS?

BUT it gives full attribution to http://en.wikipedia.org/wiki/Scraper_site?? ?

That is what paren comment says.

Re: Matt Cutts is looking for scraper sites

#27
Wikipedia's database is public and used by Google with permission. You can probably use it for your projects, too.

So this is neither scraping, nor against the rules.

Here are dumps in SQL and XML format:

http://dumps.wikimedia.org/enwiki/

Ps- Yes the original post was meant to be funny and it was; I do have a sense of humor. :)

Re: Matt Cutts is looking for scraper sites

#28
Google seems less and less connected to reality the bigger they grow.

It's a shame that the search engine market share isn't split evenly by several different engines. I think it would be beneficent both to the users and website owners. Right now everyone tries to court Google and they seem to do whatever the fuck they want.

Re: Matt Cutts is looking for scraper sites

#29
post #15

There are lots of places where Google decides to "help" me, but sometimes I just want search results. Other times, I actually like getting the curated content (e.g. search for "delta 3810"). Is there a way to disable this? EDIT: I should also note that I'm one of those who switched over to DuckDuckGo for privacy reasons, so I don't see these results as often now.

The content they do provide is often so bad it is almost embarrassing. Search for "Russia" and you get a completely useless map and a list of random facts. It may be useful to a child researching geography but for me it is just annoying.

I want content that is curated by people who actually understand the subject. I would pay for a search engine designed by someone who understands my industry. The Google algorithm only manages to grab at the low hanging fruit. I am a professional working on real stuff, I want something better than coffee shop suggestions.

Re: Matt Cutts is looking for scraper sites

#30

Scrapers lift the full content, wholesale, without attribution. You may as will just show http://images.google.com and complain that it's scraping. Or http://news.google.com . In general, do you think Wikipedia gets more traffic because Google exists, or do you think Google gets more traffic because Wikipedia exists? Meaning, which affect is larger? I'm pretty sure the answer to this is obvious. And if more scrapers…

That's one way of looking at it, on the other hand, they link to the original URL, passing traffic back to the original source. Most "scraper" sites take the content, wrap it in their own similar outer layer, and try to take ad revenue. E.g. I've seen my own StackOverflow answers copied, word for word, to a scraper site and presented under a made-up name.
Post reply on HN