Earlier quoted context omitted.
> And surprisingly, Google seems unable or unwilling to filter this stuff out. A little OT, but I find this very surprising too. In general I'm (still) a big fan of Google Search, which I think is amazing and miles above the competition. But it's true that SO clones are everywhere, often in the first five results. How can this be? How hard is it to detect that if the exact same content appeared first on SO, and then…
I suspect the problem is that Google can't fully crawl stackoverflow/github etc. There are plenty of pages which appear unindexed. I find it a lot when doing a google search for a code snippet and getting no results, only to find that exact piece of code/text is in a public github repo or bugreport. Then, the clones copy those pages, and do manage to get indexed. Which in turn means that when you search for something…
How is this possible, given their insane resources and expertise? If they can't crawl fast enough, can't they simply make a deal with Microsoft to get the public data from Github the moment it is created? I assume they have similar deals with other companies (Twitter maybe? for their firehose?).
This particular thing doesn't seem to be a tech limitation. Most likely Google doesn't care enough to fix the issue, they could probably fix it in a week if they wanted to.
At this point, nothing is going to improve until a serious competitor to Google comes along. And I don't see that happening anytime soon, so all the issues that users have with Google are likely going to stay.