Live data from Hacker News

Ask Microsoft: Are you using our personal data to train AI?

foundation.mozilla.org

51–60 of 164 posts

Re: Ask Microsoft: Are you using our personal data to train AI?

#51
post #34
post #17

Earlier quoted context omitted.

To an extent, think about vested interests here. Mozilla has little to gain by showcasing how clear a rival's new service agreement is! The AI services section seems pretty clear in terms of limiting the use cases of user content: "iv. Use of Your Content. As part of providing the AI services, Microsoft will process and store your inputs to the service as well as output from the service, for purposes of monitoring fo…

That's the only mention of AI using content. So it can be read in a few ways: 1. They will sometimes use the data for training their RLHF stuff, to "prevent harmful use" of the services. 2. The clause is exhaustive and therefore they won't use it for training, as otherwise that'd be mentioned, and are just going to log stuff for the usual monitoring purposes. This is a storm in a teacup. I don't even know why I shoul…

> I don't even know why I should care. If MS crawl some web pages I've written and AI gets slightly smarter by reading them

Crawling public web pages is a separate issue⁰ – by putting something online you aren't explicitly agreeing to any of MS's policies, at least in the eyes of the law. This is the same for anyone crawling public content not just MS.

This privacy policy covers all the content you might use MS apps and services for, i.e. where you are¹ automatically agreeing to MS's policies: OneDrive, potentially any local-only documents in Office, code in VS and other tools, perhaps anything stored on your PC running Windows.

> I don't even know why I should care.

If you don't use any MS products or services, and no products/services you do use are backed by MS's services, then you don't need to care personally. Or indeed if you do but consider everything you output or otherwise work on to be public domain. Otherwise, maybe it is something you should form an opinion on?

----

[0] time to switch my robots.txt files to “User-agent: * Disallow: /” – though it is very likely already too late for any existing content

[1] except where limited by law that you can afford to argue with MS's legal team over

Re: Ask Microsoft: Are you using our personal data to train AI?

#53
post #34
post #17

Earlier quoted context omitted.

To an extent, think about vested interests here. Mozilla has little to gain by showcasing how clear a rival's new service agreement is! The AI services section seems pretty clear in terms of limiting the use cases of user content: "iv. Use of Your Content. As part of providing the AI services, Microsoft will process and store your inputs to the service as well as output from the service, for purposes of monitoring fo…

That's the only mention of AI using content. So it can be read in a few ways: 1. They will sometimes use the data for training their RLHF stuff, to "prevent harmful use" of the services. 2. The clause is exhaustive and therefore they won't use it for training, as otherwise that'd be mentioned, and are just going to log stuff for the usual monitoring purposes. This is a storm in a teacup. I don't even know why I shoul…

No, that is true. There are multiple interpretations here. I gave the most optimistic one!

Re: Ask Microsoft: Are you using our personal data to train AI?

#54
post #44

Earlier quoted context omitted.

That paragraph says some things that they can do. It in no way says they won't use your content for AI training and any number of other things. Mozilla's point is that the whole document is sufficiently vague that they could use it to defend pretty much whatever use of your content that conceive of now or in the near future.

Why would they single out those specific uses then, if you consider express prohibitions are necessary?

To make it look, on cursory reading, like the policy is something you are comfortable to agree to. Legal theatre.

Also because those specific uses are mentioned in existing law and/or have been otherwise successfully defended. It gives their lawyers as many explicit tools as possible, before they need to argue around the implicit ones enabled by their policies & agreements being deliberately more vague elsewhere.

The point is that if they don't say that they won't, then they pretty much can if they choose to.

Re: Ask Microsoft: Are you using our personal data to train AI?

#56
post #49

So much of modern technology is a trojan horse these days. This is basically the enshittification of cloud services. Just a few years ago people would say you were being a little paranoid if you were worried about your data passing through company's servers unencrypted, but here we are now. If companies do start training models on what people consider to be private documents, then the issues we already have with AI t…

> Scientists working on papers will essentially not be able to trust that their work won't get out before they have published it.

For a long time already there's been a possibility for your work to "get out" (patented etc) before you've even finished working on it.

How did Google (and other big-tech entities) get to be the most powerful bodies on Earth?

When Google positioned themselves as "the search engine" they obtained more than a little digital privilege.

The ability to observe everything someone is searching, over a long time period, is also the ability to anticipate their moves and intentions. That's put competitors and researchers at a huge disadvantage. Researchers and competitors signal their intentions, perhaps quite unknowingly, long before they even have a clear idea of what they're doing themselves.

Of course AI only exacerbates this a million-fold. And surely what I'm saying is a decade behind the curve for anyone who is paying attention to the world.

If you have a business in tech, and are therefore a direct competitor of at least one major big-tech entity, you should not be using their services. Instead think about on-prem and local compute solutions for your most sensitive work, and relay all your search out via Tor hidden services or other mixnets for maximum diffusion .

Re: Ask Microsoft: Are you using our personal data to train AI?

#57

Earlier quoted context omitted.

> While we are talking about it... can we make ToS ilegal? What does that even mean? Laws trump Terms of service/agreements and contracts of any kind. Do they not?

I think he meant that instead of each company creating their own ToS, the government should set the standard or limitations on what a company can do. > Laws trump Terms of service/agreements and contracts of any kind. Web is not regulated by the government.

"The government"? Which one? Some governments do regulate certain aspects of the web directly.

All companies have to do is abide the local rules set by regulations, if such exist. And some very much do. Maybe not in your jurisdiction?

I think you two have things backwards

Re: Ask Microsoft: Are you using our personal data to train AI?

#58
post #39

> We had four lawyers, three privacy experts, and two campaigners look at Microsoft's new Service Agreement, which will go into effect on 30 September, and none of our experts could tell if Microsoft [will use your data] to train its AI models. * in the USA, I assume? With GDPR, if it's not a defined goal then the answer is no. In the USA, I hear things of some states having a similar law now but as a blanket stateme…

I think what these types of contracts show is that companies at Microsoft's level don't give a sh*t about national or over-regional regulation. They will use your data, and while you're busy reading 100s of pages of mumbo-jumbo they already have their models about you in place and sell access to it on a PPC basis.

And this is especially true with GDPR. Google's revenues are still growing, the advertising companies are fine, we all just have to click a few more cookie banners nowadays.

Re: Ask Microsoft: Are you using our personal data to train AI?

#59
post #17

If nine experts in privacy can't understand what Microsoft does with your data, then in my opinion a court should step in and declare it void so that Microsoft isn't allowed to use any private data until they get their act together. If it's so vague that it becomes meaningless that should default to granting no rights. Otherwise, why not publish your all-rights-granting privacy policy in Klingonian in a locked drawer…

To an extent, think about vested interests here. Mozilla has little to gain by showcasing how clear a rival's new service agreement is! The AI services section seems pretty clear in terms of limiting the use cases of user content: "iv. Use of Your Content. As part of providing the AI services, Microsoft will process and store your inputs to the service as well as output from the service, for purposes of monitoring fo…

Vested interests, yes. History, also.

For first, Mozilla doesn't do this every week. And Mozilla has a history to keep in mind general population interests for privacy and security. On the other hand, we have a corporation with a history of cheating, lying, stealing, scamming people, from fighting standards, abusing positions of power, overwriting choices going against their shareholders interests. So yeah, vested interests, but also we need to keep in mind the history of both entities

Also Mozilla didn't say "Oh we have the MS new ToS and we keep them private", they're there, get a lawyer and see if they're obvious to understand?

Re: Ask Microsoft: Are you using our personal data to train AI?

#60

If nine experts in privacy can't understand what Microsoft does with your data, then in my opinion a court should step in and declare it void so that Microsoft isn't allowed to use any private data until they get their act together. If it's so vague that it becomes meaningless that should default to granting no rights. Otherwise, why not publish your all-rights-granting privacy policy in Klingonian in a locked drawer…

[deleted]
Post reply on HN