weren't they already using repos for training?
Now, anything that gets referenced in a copilot chat is fair game
311–320 of 346 posts
weren't they already using repos for training?
Now, anything that gets referenced in a copilot chat is fair game
[flagged]
The setting isn't even visible to everyone. If you're currently in an org that manages copilot business, it's gone. I imagine it instantly opts you back in when you leave an org.
If you wouldn't use your personal email account on your work computer, I don't see why you wouldn't create a new GitHub account only for work.
No we won’t. Details here https://github.blog/news-insights/company-news/updates-to-gi... For users of Free, Pro and Pro+ Copilot, if you don’t opt out then we will start collecting usage data of Copilot for use in model training. If you are a subscriber for Business or Pro we do not train on usage. The blog post covers more details but we do not train on private repo data at rest, just interaction data with Copilot.…
Because microsuck is about to violate the law that many times
Earlier quoted context omitted.
How does that help if you don't go to the github site but just use git from the command line?
Can you use git's Copilot from the command line? If you can't, then you have nothing to opt out from.
And Copilot is integrated with IDEs. Doesn't need any interaction with the github site beyond the initial sign in...
Earlier quoted context omitted.
Algorithms and models for a proprietary trading system? My personal notes? The latex text of my phd thesis? I will go screaming and kicking and fighting into this dystopian nightmare post-privacy shithole world that so many people seem fine with. If I have to move off of every service or technology to maintain some semblance of privacy so be it.
Well, mostly I was thinking about code, and aside from the specific exceptions of trading algorithms (which I was trying to get at when I said hedge fund strategies), and now PhD theses (good point, at least if you're talking pre-publication), I'm still having trouble understanding the threat model even if AI did train on most proprietary, private business code. Can AI training on a CRUD app's code damage a business?…
Wow. This is theft. Should be illegal! It's like if I own a vault storage business and I am keeping other people's gold in my vaults and then I just take all the gold for myself and claim that the customers should have opted out of me stealing their gold but they missed the deadline...
Say some personal data leaked into training data, where can I request surgical deletion of that data from the LLM? Not only license washing is done using LLM, but also PII washing and consent ignoring is done using LLMs. How will a service provider make sure to not ever have personal data in the training data set and fix earlier mistakes pertaining to personal data? Are they not obliged to have a way of deleting one's personal data? GDPR or something?
Rather than defending this absurd decision, GitHub could instantly win back trust by admitting they f*** up and reversing it entirely. If they want to incentivise people to contribute their sources and copilot sessions, they could easily make it opt-in on a per-repository basis and provide some incentive, like an increased token quota. This is not hard.
Earlier quoted context omitted.
That comparison doesn’t hold at all. This would be equivalent to Google publishing photos of inside your home. Or, perhaps more directly, training their image-gen models on your private Google Photos.
Conceptually I think it’s a fine comparison. They’re training (with an opt out) on stuff people feel is an invasion of their privacy to make their service better.
And this is the problem: any time you're adding a new "feature" that invades users' privacy, it needs to be opt-in.