Live data from Hacker News

Federated Learning

federated.withgoogle.com

41–50 of 77 posts

Re: Federated Learning

#41

How can the data be sent in an encrypted manner that can then be useful without the server having a copy of the private keys used to encrypt the data itself?

Homomorphic Encryption

Does this actually work today? I was tangentially involved in some 'zero protocol'/zcash-related projects a few years back and the lack of ability to communicate and transfer information while being able to perform computation on it was a major drawback to most of the interesting ideas in the space.

Are they actually using this in this intended federated learning plan? If so that's a truly major innovation.

Re: Federated Learning

#42

First, I've loved that Google open sourced Tensor Flow Federated as a way to encourage the rest of the world to adopt this method of decentralized machine learning. Second, I was a bit disheartened that this concept had to be explained with a comic strip to make it accessible because I hoped the benefits were clear to everyone. Third, I read the comic strip, learned new things (secure aggregation protocol, wtf, amazi…

Popular culture is a tool for the education of the masses. Even for people who may be technically inclined, its not always evident what certain technologies really do.

I am a software engineer but mostly work on DevOps-y stuff. This was a very accessible, low-investment way for me to understand exactly what "Federated Learning" really meant.

Some of the best teachers at Univ had a way of explaining things in simple terms. This comic strip has captured that experience in a more permanent form a lot better than a textbook would.

Re: Federated Learning

#43
This is very interesting for many reasons. First we have the privacy stance, which is a tremendous step for big G. Whoever managed to push this through in the "machine" of internal office politics deserves applause. The very fact of acknowledging that users might want to control their data locally rather than rsync everything all the time is a big step—it takes us off the "give me all your data" train that we have been on for some time.

Talking about specific applications of your users' data makes a lot more sense: "If you share X with us, you're helping to build a better model Y that helps you with Z." Then the prompt "Do you want to share X?" makes a lot more sense than the current generic prompts "App V wants to access all your data W?" which doesn't tell you anything.

The anonymisation-by-aggregation aspect is interesting on it's own since it provides a practical approach we can use today and not have to wait for homomorphic encryption. There will probably still be "data leakage" but I can see how aggregation can be fundamentally better than trying to shared anonymized data by fuzzing identifiers, randomization, and binning, which are notoriously hard to pull off and suffer from de-anonymisation attacks by cross linking with other datasets.

Research-wise this could be a whole new field. Let's revisit all the ML algorithms and look at the ones that lend themselves to federated updates. Perhaps certain ML algorithms have been overlooked historically because they are not "cutting edge" but lend themselves better to distributed model updates? (I bet this is already a thing...)

The communication complexity aspects are also very interesting since it forces us to think about bandwidth needed to communicate model updates and training batching. For high-bandwidth settings we could consider training a model from scratch, for medium bandwidth you can send model updates regularly, but what would be particularly interesting to see async and VERY low bandwidth updates—like just a few MB every, exchanged once in a while when connectivity is available.

Re: Federated Learning

#44

Earlier quoted context omitted.

Looks like everyone gets the same model. So it won't be used for things like targeted ads. Cynical view is that it's only used when Google doesn't want your individual data. This will help muddy the waters in discussions about privacy.

Sure, it doesn't address that problem head on, however if they moved most stuff to a federated model, they're only a hop, skip, and a jump away from doing ad-selection on-device too, right? I think that if the model proves successful, it could end up proving out the concept enough to incentivize at least trying it out.

There are other reasons not to do things on-device. For instance, being able to pick your own ads might hinder the ability to defend against ad-viewing bots. Also, traditional ad-tech also has to deal with meeting global target spend rates (e.g. if client A pays for 1000 impressions, and client B pays for 9000 impressions, then 90% of the ads must be for client B even if every device would prefer ads for client A). Usually there are realtime databases keeping track of impression counts for all the campaigns and it would be infeasible to synchronize that with every device so that they could make local targeting decisions.

I.e. you can't necessarily make the decision locally if it's not local problem, and ad-targeting usually isn't a local problem.

Re: Federated Learning

#45

This is very interesting for many reasons. First we have the privacy stance, which is a tremendous step for big G. Whoever managed to push this through in the "machine" of internal office politics deserves applause. The very fact of acknowledging that users might want to control their data locally rather than rsync everything all the time is a big step—it takes us off the "give me all your data" train that we have be…

> "If you share X with us, you're helping to build a better model Y that helps you with Z." Then the prompt "Do you want to share X?" makes a lot more sense than the current generic prompts "App V wants to access all your data W?" which doesn't tell you anything.

That would lead to wayy too much notifications. Just like ToS, people would say yes or no blindly.

Re: Federated Learning

#46
post #29

Earlier quoted context omitted.

I think this is pretty cool from the technical perspective, but indeed I don't see the reason to be optimistic from the "privacy concerns" perspective. In fact, if this is seriously gonna be used as a "better privacy" argument (as some people seem to be already doing in this very thread), I'm calling it a PR victory for the "bad guys". First off, if you were worried about what Android was sending to Google, there's n…

It's legitimate to call the aggregate data that the central server has "non-private." For example, it's ok to publish how many thousands of cars drive on a particular highway -- that's public data. It's not ok (broadly) to publish what route John Doe drives to work everyday -- that's private data.

Absolutely not. Actually, I kind of reflected it in my concerns: it makes it easier for somebody who wants to push such a polemic (as yourself) to do so, but by no means it is universally true. I mean, it might be, if everything that your model learns is the number of cars on a highway. But there's absolutely no reason to assume it is. It might as well learn anything else about your private life.

You might be tempted to object, that it can learn nothing about your private life, since it isn't even known which device sent what data. But I think this is silly, because you forget the most important thing, which is to ask: what ever are "you"? The statement in question depends on definition of that, because if "you" means "your IP address", it might be true. But I never really was afraid Google is learning something about what is sent from this IP, because I don't believe the give a fuck.

It is much less apparent they don't learn anything about "John Doe" simply given the fact that the learning is "federated". But then again, this doesn't really bother me much, because I don't believe they care.

I think, much more probable definitions of "you" that might concern them, are a lot more dangerous, because they are a lot more "real" than your made-up (even if at birth) name which you probably share with 100 more people around the world anyway. Like, for example, "the guy, who every day makes a trip from Baker street 12 to Sesame street 30" (that's "you") and that he really likes Coke. And this is just the most simplistic one, there are infinite tuples of parameters that would define a specific person in much more meaningful way, than the IP or a name.

But, actually, I don't really believe they would try to identify you like that either. Well, they might, I just don't really believe they care about "you" in a sense that might be meaningful to you. What some entity like Google is likely to mean by "you" is probably several organisms, so you might feel like whatever they learn about "you" is definitely not private data. But I don't think this is less scary, quite the opposite, facts like "boys 13-15 y.o. that listen Tokyo Hotel and drink Sprite are likely to try heroine if recommended to watch RocknRolla" are much more powerful and useful (this is obviously a made up example, which you may replace by anything seemingly less dramatic, like "person, who buys A, B and C will also want to buy D"). Anyway, what really makes up a model that can control the financial markets and mood of the people isn't about your petty definition of "you", but a much colder, more meaningful one.

And when I'm worried Google learns something about "me", this is what I'm really worried about. And the fact that such a definition of "private data" wouldn't hold up in court because it "isn't even about a specific person" makes it only so much worse.

Re: Federated Learning

#47

Google mentioned at I/O that speech recognition will soon (this summer?) be performed locally on Android devices, with no voice data being sent to Google, because they have been able to reduce the size of the model dramatically. Is that related to federated learning? Paper: https://arxiv.org/abs/1811.06621

No, federated learning is about training on the device not running prediction on the device. Training a speech model on the device would be hard, because there is no labeled data. We don't know what the user said.

Re: Federated Learning

#48
post #29

Earlier quoted context omitted.

I think this is pretty cool from the technical perspective, but indeed I don't see the reason to be optimistic from the "privacy concerns" perspective. In fact, if this is seriously gonna be used as a "better privacy" argument (as some people seem to be already doing in this very thread), I'm calling it a PR victory for the "bad guys". First off, if you were worried about what Android was sending to Google, there's n…

This is not the way it works at all (see Secure Aggregation). There are a number of techniques out there that permit privacy safe services and learning, like Differential Privacy, Federated Learning, and there's even Deep Neural Nets using homomorphic computing (e.g. https://arxiv.org/abs/1711.05189 ) How about we lay off the conspiracy theories every time any new paper is published.

"Secure aggregation" doesn't mean a thing, because I would argue, as I did in another comment already, that "you" are not "your device".

Re: Federated Learning

#49
post #27
post #18

Earlier quoted context omitted.

Well the cynical view would be. 1) This still lets you have personalized models, just trained on more than 1 user, thats fine at google's scale anyway 2) Their competitors (FB, AMZN) dont have the edge compute (Android) to do this, and to a lesser degree don't have the ML stack (however Android implements this at the API level will be very Tensorflow focused) 3) Now google can push for privacy regulations that preven…

If #3 happens I'd be shocked (and pleased).

They would be absolutely compelled to by market forces. The entire ad-tech industry would be wiped out if something like this becomes more common. You would need an army of highly trained machine learning engineers to build complex systems that would work great but still be efficient. Guess who is the only company which has that army?

The cynic in me sees this as a great play by the Googz to cut the Amazon-adtech venture in the bud, and to establish and maintain dominance over the adtech business, and advertising in general.

Re: Federated Learning

#50
post #45

This is very interesting for many reasons. First we have the privacy stance, which is a tremendous step for big G. Whoever managed to push this through in the "machine" of internal office politics deserves applause. The very fact of acknowledging that users might want to control their data locally rather than rsync everything all the time is a big step—it takes us off the "give me all your data" train that we have be…

> "If you share X with us, you're helping to build a better model Y that helps you with Z." Then the prompt "Do you want to share X?" makes a lot more sense than the current generic prompts "App V wants to access all your data W?" which doesn't tell you anything. That would lead to wayy too much notifications. Just like ToS, people would say yes or no blindly.

Well, I really would like to say "no" to all of them, but somehow I don't expect to be given an option.
Post reply on HN