What's the best way to poison my repos to sabotage LLM training? Asking for a friend.
Which doesn't answer your question at all, but it is the metric they'll pay attention to. And it is the the thing that actually addresses the underlying problem.
281–290 of 346 posts
What's the best way to poison my repos to sabotage LLM training? Asking for a friend.
Which doesn't answer your question at all, but it is the metric they'll pay attention to. And it is the the thing that actually addresses the underlying problem.
Earlier quoted context omitted.
What are you referring to? I set this to "Disabled" months/years ago and it's retained the disabled setting.
So? You guarantee that this setting is durable and will never revert? Or you guarantee that no client-side bug on that page will not override the setting with null value when you click save on something else? Please.
The chip on your shoulder doesn't make for productive conversation here.
Jokes on them, my private repos are total dog dookie. If nobody but me can see the code then I don't have to worry about style, structure, comments, or any other best practices. You don't want an LLM trained on my private repos. Trust me.
Poisoning LLMs is an interesting path of resistance.
Earlier quoted context omitted.
I don't like when people make sarcastic remarks and sign off in a way that indicates it was sarcasm. It kills it for me. Lol. Like using that /s or using that smiling emoji sign you used. A good joke would land even if some other people miss it because of the text format. "Microsoft would never do this" would have landed for me.
In Soviet Russia, Microsoft is the shit!
No we won’t. Details here https://github.blog/news-insights/company-news/updates-to-gi... For users of Free, Pro and Pro+ Copilot, if you don’t opt out then we will start collecting usage data of Copilot for use in model training. If you are a subscriber for Business or Pro we do not train on usage. The blog post covers more details but we do not train on private repo data at rest, just interaction data with Copilot.…
No we won’t. Details here https://github.blog/news-insights/company-news/updates-to-gi... For users of Free, Pro and Pro+ Copilot, if you don’t opt out then we will start collecting usage data of Copilot for use in model training. If you are a subscriber for Business or Pro we do not train on usage. The blog post covers more details but we do not train on private repo data at rest, just interaction data with Copilot.…
""" Allow GitHub to use my data for AI model training
Allow GitHub to collect and use my Inputs, Outputs, and associated context to train and improve AI models. Read more in the Privacy Statement. """
If the reality is less scary than how it sounds, then the wording needs to be less scary-sounding. It may be that GitHub isn't training models on private repos, but the language certainly suggests that it is. The feedback we're seeing in this post is proof enough of that.
Finally, I read the Privacy Statement, and it's unclear what the applicable language is. "Inputs," "Outputs," and "Associated Context" are terms of art that have no matching definitions in the Statement. (The terms "Outputs" and "Associated Context" don't even appear in the Statement at all. Not even "train.") As an attorney I find this completely baffling.
Earlier quoted context omitted.
Microsoft would never do this (-:
I don't like when people make sarcastic remarks and sign off in a way that indicates it was sarcasm. It kills it for me. Lol. Like using that /s or using that smiling emoji sign you used. A good joke would land even if some other people miss it because of the text format. "Microsoft would never do this" would have landed for me.
I love falling into a rabbit hole looking at people’s projects