Live data from Hacker News

Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

news.ycombinator.com

41–50 of 67 posts

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#41
When I look at w3facility.org, it seems like Googles algorithms do not properly handle the case when a scrape-site provides a source-link to the original content.

Google recommends the latter to protect against duplicate content penalties when you use some external content to enrich your site (for example a short section from Wikipedia, imdb actor info, etc).

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#42
I see this happen a lot also with the MSDN forums. The interesting thing there is that some of the mirror sites are still carrying topics that have been deleted or otherwise disappeared from the real MSDN forums. More than once, in the obscure subset of Microsoft tooling that I work in, the only hits that are still alive are on somewhat suspicious looking .ru sites, so in some sense, I am glad these sites do exist - otherwise, I'd be completely SOL trying to figure out why the badly-documented API I'm relying on is barfing up an opaque HRESULT.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#43
post #13

Google doesn't have a good way to establish provenance, and has trouble distinguishing copies from originals. It's a common complaint of blog operators that some bigger blog copied their stuff and got a higher ranking on Google. Google could check when it saw something, but that won't work against fast scrapers. For that, you need trusted timestamps. One solution to this would be to have a few time-stamping services.…

Timestamping sounds like a good idea. We can easily do it now in a decentralized way without depending on any 3rd party service e.g. https://github.com/chainpoint/chainpoint

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#44
post #21

Earlier quoted context omitted.

We use that license because it protects the content from us. No matter who comes along to run Stack Overflow in the future, Stack Overflow can't do something like put up a paywall and lock it up. Someone else'll just be able to host a copy.

As an end user, I don't care about that at all. If StackOverflow does something like put up a paywall and lock it up, then it will die off and some other site will arise that will replace it, just like some other sites that came before StackOverflow which put up a paywall and faded away. Also, StackOverflow can always change the policy if they want (which probably won't happen for the reason I mentioned), so the lice…

Basically, I'm not willing to supply useful data to a site which then profits off of it without also making it open.

So if they weren't open then they wouldn't be getting my answers (or a bunch of other people's).

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#45
post #21

Earlier quoted context omitted.

We use that license because it protects the content from us. No matter who comes along to run Stack Overflow in the future, Stack Overflow can't do something like put up a paywall and lock it up. Someone else'll just be able to host a copy.

As an end user, I don't care about that at all. If StackOverflow does something like put up a paywall and lock it up, then it will die off and some other site will arise that will replace it, just like some other sites that came before StackOverflow which put up a paywall and faded away. Also, StackOverflow can always change the policy if they want (which probably won't happen for the reason I mentioned), so the lice…

A new site may come to replace it, but what of the 10MM+ questions and scads more answers created by the Stack Overflow community over the last 7 years? Without the CC license all that hard work gets lost in the case SO goes rogue.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#46
post #19
post #17

Earlier quoted context omitted.

Since establishing provenance is such a big problem for Google, perhaps it might be a good idea for Google to offer a time-stamping service itself?

It's not directly a problem for Google — it's primarily a problem for sites that create original content.

It's a problem that would disappear if content creators were compensated regardless of serving point.

We don't have that. We could.

Universal content syndication + broadband tax.

https://www.reddit.com/r/dredmorbius/comments/1uotb3/a_modes...

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#47
post #21

Earlier quoted context omitted.

We use that license because it protects the content from us. No matter who comes along to run Stack Overflow in the future, Stack Overflow can't do something like put up a paywall and lock it up. Someone else'll just be able to host a copy.

As an end user, I don't care about that at all. If StackOverflow does something like put up a paywall and lock it up, then it will die off and some other site will arise that will replace it, just like some other sites that came before StackOverflow which put up a paywall and faded away. Also, StackOverflow can always change the policy if they want (which probably won't happen for the reason I mentioned), so the lice…

The license means that even if SO changes its policy and puts up a paywall, a clone can come up, use the knowledge already present on SO that was scraped.

This means that even if SO puts a paywall, the knowledge gathered there is not lost. And this, IMO, is a very, VERY important aspect of StackOverflow.

As far as search goes, I didn't even know SO had a search function. I see nothing wrong with letting google handle it. Adding a search engine to your website is hard and takes time.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#48
post #39
post #21

Earlier quoted context omitted.

As an end user, I don't care about that at all. If StackOverflow does something like put up a paywall and lock it up, then it will die off and some other site will arise that will replace it, just like some other sites that came before StackOverflow which put up a paywall and faded away. Also, StackOverflow can always change the policy if they want (which probably won't happen for the reason I mentioned), so the lice…

Wow i'm getting downvoted like crazy. I don't think I said anything that's not factual. At least explain why you think I am wrong if you're gonna downvote. To be clear, I love Stackoverflow and I don't know what the world would have been like if it wasn't around, but I do think there are things that are broken and I just mentioned them. Am I supposed to keep quiet because that's how it's been?

"Please resist commenting about being downvoted. It never does any good, and it makes boring reading."

https://news.ycombinator.com/newsguidelines.html

(NB: I didn't, and tend to agree it's excessive here.)

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#49
post #21

Earlier quoted context omitted.

As an end user, I don't care about that at all. If StackOverflow does something like put up a paywall and lock it up, then it will die off and some other site will arise that will replace it, just like some other sites that came before StackOverflow which put up a paywall and faded away. Also, StackOverflow can always change the policy if they want (which probably won't happen for the reason I mentioned), so the lice…

A new site may come to replace it, but what of the 10MM+ questions and scads more answers created by the Stack Overflow community over the last 7 years? Without the CC license all that hard work gets lost in the case SO goes rogue.

To be fair, useful life of many of those answers is likely only a few years at best. Though I generally agree: locking up contributed content is exceptionally poor Internet citizenship (Quora, Scribd).

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#50
post #17

Earlier quoted context omitted.

Since establishing provenance is such a big problem for Google, perhaps it might be a good idea for Google to offer a time-stamping service itself?

The problem is how one prevents the bad-actors from taking advantage of such a service.

Define "bad actor".

That's not a "they aren't", but a request for "what do you consider to be acting badly?"

Conversely, is content scraping-and-publication possible by good actors?

How?

Post reply on HN