Live data from Hacker News

Microsoft’s purchase of GitHub leaves some scientists uneasy

nature.com

51–60 of 148 posts

Re: Microsoft’s purchase of GitHub leaves some scientists uneasy

#51

Earlier quoted context omitted.

It was a bad idea in the first place, and one of the dangers was that this one failure point could end up controlled by a company like Microsoft. You can say that it was a bad idea for people to store all of their things in a warehouse with no sprinklers or smoke detectors, but after it's been done, it's pretty silly to scold people for not continuing to store things in a warehouse that is currently on fire.

> for not continuing to store things in a warehouse that is currently on fire. That metaphor is rather silly. Microsoft's purchase of github didn't set the metaphorical warehouse on fire. At most, it only raised attention to the fire hazard that was present since the warehouse was created. What next? People will complain about the privacy risks presented in services such as Dropbox only after some major corportion bu…

I think it's blindingly naive to think that, on any meaningful timeline, a company like GitHub wouldn't be mining their userbase for trends and capitalising on them by selling it to third parties... There's nothing shady about using corporate information and user metadata to provide value to others, but you have to imagine that MS owning that data set is highly similar to MS buying it from GitHub.

Up to the minute analysis of a meaningful percentage of the development ecosphere is highly valuable. Selling reports, analysis, or monitoring is a natural expansion of GitHubs business model.

Maybe having this dataset go into the same warehouse as LinkedIn strikes some as scary, but I think one actor who is less reliant on direct revenue will have less incentive to push boundaries and spread that info to as many other actors as possible... So whatever privacy issue people feel has arisen here, it's likely just refined itself a little to be much narrower but only a little deeper.

Fundamentally, if you don't trust MS with your goodies you shouldn't trust GitHub with your data at all. GitHub Enterprise might be a better solution ;)

Re: Microsoft’s purchase of GitHub leaves some scientists uneasy

#52
post #40

Earlier quoted context omitted.

they intend to keep VSTS and github as separate product Still, consider Codeplex, which isn’t around anymore. I see the two merging over time, say in 5 years, and I expect the migration to be fairly painless, so I am not too worried personally, but I can see how some people might be.

Unlikely, they're targeting different markets. There is a place for both of them in Microsoft's portfolio.

MS-GitHub Enterprise would seem to address a lot of where TFS otherwise lands. Once TFS incorporates GitHub workflows and infrastructure I think anything off of that core is gonna struggle mightily.

I mean... as it stands TFS struggles with mindshare, features, and engagement. While the name and marketing will persist, I think the underlying technical reality will be the continued cannibalisation of that offering by the superior model of git and its cross-language appeal.

Re: Microsoft’s purchase of GitHub leaves some scientists uneasy

#53
post #48

To survive, GItHub must provide the way to encrypt private repositories so that they can make sure that files cannot be read by github employees. Would that be possible?

Microsoft is the number two cloud vendor. When they violate the privacy of their paying customers, they are dead.

And no, telemetry is about privacy but not the same as reading the data/code itself.

Re: Microsoft’s purchase of GitHub leaves some scientists uneasy

#54
post #38

Seeing this published on Nature's website is quite thick with irony. They have profited greatly from locking up access to academic papers in their most prestigious journals. The article includes the following tweet: “Open Science is not compatible with one corporation owning the platform used to collaborate on code. I hope that expert coders in #openscience have a viable alternative to #github,” tweeted Tom Johnstone…

Nature is also not an Open Access or Open Science journal at all, and furthermore, does not require authors to publish all of their data for replication purposes.

It's incredibly hypocritical. If people cared about Open Science, they'd publish in PeerJ or PLOS One, but they care a lot more about saying they have pub credits in Nature.

Re: Microsoft’s purchase of GitHub leaves some scientists uneasy

#55
post #51

Earlier quoted context omitted.

> for not continuing to store things in a warehouse that is currently on fire. That metaphor is rather silly. Microsoft's purchase of github didn't set the metaphorical warehouse on fire. At most, it only raised attention to the fire hazard that was present since the warehouse was created. What next? People will complain about the privacy risks presented in services such as Dropbox only after some major corportion bu…

I think it's blindingly naive to think that, on any meaningful timeline, a company like GitHub wouldn't be mining their userbase for trends and capitalising on them by selling it to third parties... There's nothing shady about using corporate information and user metadata to provide value to others, but you have to imagine that MS owning that data set is highly similar to MS buying it from GitHub. Up to the minute an…

Except that selling data to third parties has never been Microsoft’s MO. It’s always been about getting fat checks from the worlds largest companies’ IT departments.

Re: Microsoft’s purchase of GitHub leaves some scientists uneasy

#56

To be honest, I don't think GitHub is to be blamed for how scientists use it. AFAIK there are data repositories more suitable for the purpose of sharing data sets and metadata, and out of the hands of big corporations (but may subject to big governments), e.g., Zenodo (funded by CERN) and DataONE (funded by NSF). Zenodo can even generate a DOI for your data set, which GitHub does not do. That being said, I think this…

Microsoft Research is one of the largest supporters of academic computer science research in the world, and is a non-trivial funding source for many non-CS academic fields as well.

Re: Microsoft’s purchase of GitHub leaves some scientists uneasy

#57
post #50
post #38

Seeing this published on Nature's website is quite thick with irony. They have profited greatly from locking up access to academic papers in their most prestigious journals. The article includes the following tweet: “Open Science is not compatible with one corporation owning the platform used to collaborate on code. I hope that expert coders in #openscience have a viable alternative to #github,” tweeted Tom Johnstone…

See I hope that "expert" coders understand that 'git push' works even on non-MS controlled endpoints, and how to setup a damned wiki if need be... GitHub is stupid valuable for MS, and will make their dev tools offerings much stronger (TFS is a dog, bringing GitHub Enterprise into MS Enterprise support & sales cycle is gonna make a loooot of scratch). But the value of the GitHub community persists only so long as the…

You can already use GitHub in visual studio. MS didn't have to buy GitHub to integrate with it.

Pretty much the best that can happen is GitHub will now get worse as MS puts in functionality to do with their tools. The worst is they really fuck it up like they did with Skype.

Re: Microsoft’s purchase of GitHub leaves some scientists uneasy

#58

To be honest, I don't think GitHub is to be blamed for how scientists use it. AFAIK there are data repositories more suitable for the purpose of sharing data sets and metadata, and out of the hands of big corporations (but may subject to big governments), e.g., Zenodo (funded by CERN) and DataONE (funded by NSF). Zenodo can even generate a DOI for your data set, which GitHub does not do. That being said, I think this…

“Why don’t scientists use modern tools like GitHub, instead of these crap bespoke ones? They shouldn’t waste their time duplicating effort.” “Why are scientists using inappropriate tools? These are for commercial use only!” In practice, GitHub works fine for collaborative code projects and sharing data over the next few years. If you’re looking to store large amounts of data over long periods, you’re right. If GitHub…

Not cool to insert fake quotes as if you are quoting the person you’re responding to.

Re: Microsoft’s purchase of GitHub leaves some scientists uneasy

#59
post #25

This article seems to have a very focused interest on data in GitHub repositories, as opposed to source code. I get that the article is aimed at scientists, but I don’t see the problem here: if Microsoft takes down your dataset just move it somewhere else. You’re not tied down with pull requests or comments like code repositories are.

The problem is the stable of the pointer. From a paper of mine: "The code and simulation results are available online at https://github.com/mylabname/ project." That's in print . In another paper, I expressly cite a GitHub repository as the source of the data used in the analysis. Pointing to data in papers is the way most of the scientists I know use GitHub - because it's relatively stable, not tied to an institutio…

The problem is not any particular hosting service such as GitHub, it's that the pointer in your paper ss relying on a single point of failure.

There was a bit of drama several years ago when Megaupload was seized and shut down; various small/free projects lost access to the only copy of some of their files. Like your paper, important documents had evolved in forums, which linked to the file hosting service for files that could not be uploaded to the forum. A few projects were the canonical documentation for something that the original author had abandoned, the first result in Google couldn't be updated creating the same pointer problem as your reference in a paper.

At the time, a lot of people talked about finding a "replacement file hosting service" in the same way people currently talk about finding a replacement for GitHub. Moving to a different service is still a single point of failure. Instead, when you want to preserve access to data in the long term, you need to assume any single service might fail and build in redundancy.

Instead of saying, "[things] are available online at [URL]", you should include in the paper something like:

    The code and simulation results are
    available as an archive named:
           foo_project-2018-06-16.zip
    The file has the following checksums:
        MD5:  1271ed5ef305aadabc605b1609e24c52
        SHA1: ab69db8315af7de6e673a6ddf128d415157a7c3f
        (...more...)
    The file was originally hosted at:
        $GITHUB_URL
        $GITLAB_URL
        ${OTHER_HOSTING_SERVICE_URLS[@]}
        $INTERNET_ARCHIVE_URL
        $AUTHOR_UNIVERSITY_URL
        $COLLAB_UNIVERSITY_URL
With tools like git (or rsync, etc), making multiple copies of a project is very easy. Redundancy protects against some risks, but including checksums (and any other relevant metadata) makes content addressable searching possible. Even if all of the URLs in your paper eventually become defunct, someone reading the paper in the future may be able to find your data by searching for the file's hash.

The hosting service isn't the (primary) problem; the paper needs to include a pointer that is more robust than a single reference to a single path on a single server.

Re: Microsoft’s purchase of GitHub leaves some scientists uneasy

#60
post #31
post #25

Earlier quoted context omitted.

The problem is the stable of the pointer. From a paper of mine: "The code and simulation results are available online at https://github.com/mylabname/ project." That's in print . In another paper, I expressly cite a GitHub repository as the source of the data used in the analysis. Pointing to data in papers is the way most of the scientists I know use GitHub - because it's relatively stable, not tied to an institutio…

[disclaimer work for related company - Digital Science] It's really worth getting those things into a system that will give you a doi and participate in an archive so that content should never be lost. Figshare is one, zenodo another (I think but can't find on a phone that the data is archived). A doi, if versioned, can ensure people see the actual version you used. It can be pointed somewhere else if the provider go…

I've been consistently unimpressed with the offerings in that space.
Post reply on HN