Live data from Hacker News

Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

news.ycombinator.com

31–40 of 67 posts

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#31
post #17

Earlier quoted context omitted.

Since establishing provenance is such a big problem for Google, perhaps it might be a good idea for Google to offer a time-stamping service itself?

Google offers this service already: http://www.labnol.org/internet/fat-pings-for-content-scraper... https://en.wikipedia.org/wiki/PubSubHubbub https://pubsubhubbub.appspot.com/ Most of the big hosted publishing platforms like WordPress and Blogger already use it, but it's pretty common for sites that built their own codebase not to. (This was one of my interests at Google, and I had both a 20% project [unreleased] an…

That's not a timestamping service; that's just a way to quickly notify Google when you add a page to your blog.

What if your blog doesn't use Pubsubhubbub, but mine does, and I copy your blog post? Google will see my post first, but that doesn't prove I wrote it, so Google cannot use that to make ranking decisions.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#32
post #31

Earlier quoted context omitted.

Google offers this service already: http://www.labnol.org/internet/fat-pings-for-content-scraper... https://en.wikipedia.org/wiki/PubSubHubbub https://pubsubhubbub.appspot.com/ Most of the big hosted publishing platforms like WordPress and Blogger already use it, but it's pretty common for sites that built their own codebase not to. (This was one of my interests at Google, and I had both a 20% project [unreleased] an…

That's not a timestamping service; that's just a way to quickly notify Google when you add a page to your blog. What if your blog doesn't use Pubsubhubbub, but mine does, and I copy your blog post? Google will see my post first, but that doesn't prove I wrote it, so Google cannot use that to make ranking decisions.

I don't see how pubsubhubbub differs from the originally proposed timestamping service. For any timestamping service, it will still be the case that if you don't use it, and somebody else scrapes you and does use it, their timestamp will be better than yours.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#33
post #13

Google doesn't have a good way to establish provenance, and has trouble distinguishing copies from originals. It's a common complaint of blog operators that some bigger blog copied their stuff and got a higher ranking on Google. Google could check when it saw something, but that won't work against fast scrapers. For that, you need trusted timestamps. One solution to this would be to have a few time-stamping services.…

Isn't part of the core of pagerank that certain domains are more trustworthy than others? Given that this is supposedly true, wouldn't the same content on stackoverflow out rank a random new domain with the same content

Exactly. In the SEO world, the common wisdom is you will get dinged in ranking or even deindexed if you publish duplicate content. Period. There is an entire sub-industry in SEO that's all about content spinning. I'm always kinda amazed that these sites show up and I usually get them for about 50%+ of my tech searches.

What I think ultimately happens is Google is notified and these sites are eventually deindexed. But this still leads me to really question the whole duplicate content theory sometimes. If Google was so tight about this, these sites would never show up in the first place.

It could also be that the tech terms searched for are not as competitive as, say, medical or diet or celeb terms. Maybe Google's just grabbing at any kind of relevancy at this point, duplicate or not.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#34
Sort of related, but in the last year it seems to me that Google's results in general have been increasingly getting worse. Now it seems that at least two or three spam results are always present in the first page of results. Most of these results also happen to be duplicate content of bigger sites though. I have tried reporting on numerous occasions to Google, but just usual "we will investigate" response and never a word back.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#35
post #13

Google doesn't have a good way to establish provenance, and has trouble distinguishing copies from originals. It's a common complaint of blog operators that some bigger blog copied their stuff and got a higher ranking on Google. Google could check when it saw something, but that won't work against fast scrapers. For that, you need trusted timestamps. One solution to this would be to have a few time-stamping services.…

Then if the original didn't use timestamping, and the clones did, they'd always win over the original.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#36
post #17
post #13

Google doesn't have a good way to establish provenance, and has trouble distinguishing copies from originals. It's a common complaint of blog operators that some bigger blog copied their stuff and got a higher ranking on Google. Google could check when it saw something, but that won't work against fast scrapers. For that, you need trusted timestamps. One solution to this would be to have a few time-stamping services.…

Since establishing provenance is such a big problem for Google, perhaps it might be a good idea for Google to offer a time-stamping service itself?

The problem is how one prevents the bad-actors from taking advantage of such a service.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#37

Earlier quoted context omitted.

Isn't part of the core of pagerank that certain domains are more trustworthy than others? Given that this is supposedly true, wouldn't the same content on stackoverflow out rank a random new domain with the same content

Exactly. In the SEO world, the common wisdom is you will get dinged in ranking or even deindexed if you publish duplicate content. Period. There is an entire sub-industry in SEO that's all about content spinning. I'm always kinda amazed that these sites show up and I usually get them for about 50%+ of my tech searches. What I think ultimately happens is Google is notified and these sites are eventually deindexed. But…

This is really surprising. Given that they have so much data, there should at least be several thousand domains that they treat with some special regard when content that replicates them shows up.

Then again, I'm sure that there are plenty of top 1000 domains that would use this to their advantage too, oh well.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#38
post #13

Google doesn't have a good way to establish provenance, and has trouble distinguishing copies from originals. It's a common complaint of blog operators that some bigger blog copied their stuff and got a higher ranking on Google. Google could check when it saw something, but that won't work against fast scrapers. For that, you need trusted timestamps. One solution to this would be to have a few time-stamping services.…

Isn't part of the core of pagerank that certain domains are more trustworthy than others? Given that this is supposedly true, wouldn't the same content on stackoverflow out rank a random new domain with the same content

pagerank works both at the domain and page level, originally with the emphasis on page level(hence pagerank not domainrank).

While it would be hard for these sites to match the cumulative reach of SO, they will have an easier time getting specific pages to rank highly. This can also be abused with a system where most pages link to the .1% of pages that should be emphasized. In this fashion, smaller sites can throw their weight around and get specific pages to rank more highly than bigger sites.

Trying to solve this issue on a purely technical level turns out to be a lot more complicated than it would seem. It is made much worse by how damaging false-positives are for a search engine (Google censored me!). So the result is that this often only gets resolved by manual actions rather than automatic penalties.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#39
post #21

Earlier quoted context omitted.

We use that license because it protects the content from us. No matter who comes along to run Stack Overflow in the future, Stack Overflow can't do something like put up a paywall and lock it up. Someone else'll just be able to host a copy.

As an end user, I don't care about that at all. If StackOverflow does something like put up a paywall and lock it up, then it will die off and some other site will arise that will replace it, just like some other sites that came before StackOverflow which put up a paywall and faded away. Also, StackOverflow can always change the policy if they want (which probably won't happen for the reason I mentioned), so the lice…

Wow i'm getting downvoted like crazy. I don't think I said anything that's not factual. At least explain why you think I am wrong if you're gonna downvote. To be clear, I love Stackoverflow and I don't know what the world would have been like if it wasn't around, but I do think there are things that are broken and I just mentioned them. Am I supposed to keep quiet because that's how it's been?

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#40
This is common across numerous contexts. I've found what appear to be bots Tweeting my reddit and HN posts (I don't mind), several Reddit clones of various levels of sniffitude, Google+ content harvesting, and some Diaspora syndication (to be expected), again, of various levels of sniffitude.

To the extent that this simply distributes data around, doesn't claim it for its own, and credits source, I'm OK with this. Better even if it follows site-specific licensing. Among my visions for the Web would be content syndication where such schemes would actually directly benefit authors and creators, regardless of where their content is served.

Post reply on HN