Live data from Hacker News

Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

news.ycombinator.com

61–67 of 67 posts

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#61
post #6

I don't understand why StackOverflow allows this, yeah creative commons is cool and all but it IS NOT COOL for actual users. I get so annoyed every time i search for something on Google and it leads to an its clone site. It's not like StackOverflow has better search than Google (which is ridiculous). I still have to search on Google if I want quality search result instead of StackOverflow. With more power comes more…

I once contributed to a localized version of SO whose licence was not open. The startup backing it eventually shutdown and the site disappeared. All time invested in those answers is now irrevocably lost.

Non-open content licence is exactly what made me stop contributing to sites such as Wikimapia and start contributing to OpenStreetMap instead.

Once you get burned a few times, you learn it's better to have to put up with a few clone sites than risk losing all the data, or having to pay to read what you wrote yourself a few months ago.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#62
post #61
post #6

I don't understand why StackOverflow allows this, yeah creative commons is cool and all but it IS NOT COOL for actual users. I get so annoyed every time i search for something on Google and it leads to an its clone site. It's not like StackOverflow has better search than Google (which is ridiculous). I still have to search on Google if I want quality search result instead of StackOverflow. With more power comes more…

I once contributed to a localized version of SO whose licence was not open. The startup backing it eventually shutdown and the site disappeared. All time invested in those answers is now irrevocably lost. Non-open content licence is exactly what made me stop contributing to sites such as Wikimapia and start contributing to OpenStreetMap instead. Once you get burned a few times, you learn it's better to have to put up…

Thanks for the answer. I just want to ask one more question. Don't you think StackOverflow is already too large to disappear suddenly like those startups you mentioned, and smart enough to NOT go the way of ExpertsExchange? (I think the founder explicitly mentioned StackOverflow wanted to be an anti ExpertsExchange when they launched. Also I don't think the company would be stupid enough to alienate their users which will lead to demise regardless of how much content they already have)

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#63

Earlier quoted context omitted.

Isn't part of the core of pagerank that certain domains are more trustworthy than others? Given that this is supposedly true, wouldn't the same content on stackoverflow out rank a random new domain with the same content

pagerank works both at the domain and page level, originally with the emphasis on page level(hence pagerank not domainrank). While it would be hard for these sites to match the cumulative reach of SO, they will have an easier time getting specific pages to rank highly. This can also be abused with a system where most pages link to the .1% of pages that should be emphasized. In this fashion, smaller sites can throw th…

Actually it's pagerank for Larry Page

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#64

Earlier quoted context omitted.

pagerank works both at the domain and page level, originally with the emphasis on page level(hence pagerank not domainrank). While it would be hard for these sites to match the cumulative reach of SO, they will have an easier time getting specific pages to rank highly. This can also be abused with a system where most pages link to the .1% of pages that should be emphasized. In this fashion, smaller sites can throw th…

Actually it's pagerank for Larry Page

hence his name is not Larry Domain.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#65
post #17

Earlier quoted context omitted.

Since establishing provenance is such a big problem for Google, perhaps it might be a good idea for Google to offer a time-stamping service itself?

Exactly! It seems to me Google could easily provide a proof-of-authorship API, especially for text. (You could publish a hash in some feed in case you don't fully trust Google in turn.) I'm not a big fan of conspiracies and such, but by now I'm pretty sure Google has some conflict of interest given that it hasn't offered such a service already.

Problem as others pointed out is that no matter how you design the system, even with Google a completely as trusted party in this, proof-of-authorship, timestamping-service, it can only work if every original content creator (or platform, like SO) gets on this train.

Say you're a content creator that maybe doesn't want to use this service, simply hasn't heard about it yet, or accidentally publishes a whole archive of original articles they forgot to sign with the proof-of-authorship/timestamping API just before they get scraped by an ill-intentioned content farmer, that quickly uses the API to sign/stamp the articles to themselves before the content creator can. They don't even need to publish them right away, they can drip them out over years and have the proof-of-authorship signature to "prove" their authorship.

If we would actually trust this system, it means that the real original content creator is shit out of luck. There's no way they can prove their authorship in a way that distinguishes them from a scraper, that part remains the same with or without this system--but what would be new is that the scraping non-author now has some sort of extra claim of authenticity over the actual original content creator.

I see no way around this. Unless you devise a way so that all content written (anywhere, any time, on any medium) is immediately signed with a proof-of-authorship, such a system will be giving power to bad actors to claim authorship of any content that happens to slip through without being signed.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#66
post #6

I don't understand why StackOverflow allows this, yeah creative commons is cool and all but it IS NOT COOL for actual users. I get so annoyed every time i search for something on Google and it leads to an its clone site. It's not like StackOverflow has better search than Google (which is ridiculous). I still have to search on Google if I want quality search result instead of StackOverflow. With more power comes more…

We use that license because it protects the content from us. No matter who comes along to run Stack Overflow in the future, Stack Overflow can't do something like put up a paywall and lock it up. Someone else'll just be able to host a copy.

Well, that was a bad decision with regards to UX because now I have to dig through 20 look-alike sites that don't link back to SO and you're not doing shit about it. Yes, SO comes up first when I search for the exact question but I'm sure you're aware how useless that is. If I knew the exact question to ask, I probably wouldn't be searching around Google.

My time is the most valuable thing to me and with this decision, you've essentially costed me a lot of time and I already paid my way by adding content to your site. I would have much rather seen you enter a simple legal agreement where the content was placed in escrow.

Re: Ask HN: Has anyone else noticed Stack Overflow clones in Google search results?

#67

Earlier quoted context omitted.

pagerank works both at the domain and page level, originally with the emphasis on page level(hence pagerank not domainrank). While it would be hard for these sites to match the cumulative reach of SO, they will have an easier time getting specific pages to rank highly. This can also be abused with a system where most pages link to the .1% of pages that should be emphasized. In this fashion, smaller sites can throw th…

Actually it's pagerank for Larry Page

Actually it's not. (In case anyone might think you were being serious.)
Post reply on HN