Live data from Hacker News

Google is making private AI practical with homomorphic encryption

blog.google

261–270 of 305 posts

Re: Google is making private AI practical with homomorphic encryption

#261
post #225

Earlier quoted context omitted.

Then the app can make a bog-standard encrypted-at-rest backup to somewhere and make all the computations on the device on the cleartext data. I don't see the need to do computations on the encrypted data here, which is what FHE would provide in addition to traditional encryption. > and don't have access to the account anymore. This would be trouble with or without FHE. Even if the backend wouldn't need to decrypt the…

Okay, so the platform becomes valuable to its users when it's able to suggest things like "based on millions of users, people with cycles like yours typically ovulate around day 16." In order to do the data mining in order to make those kinds of claims, traditionally you'd need to have access to the data. As you point out, encrypted-at-rest is solved. But what about when it's not at rest? In-use and in-transit is whe…

If you can trust the app/platform developers to send encrypted anonymised personal data, then how can you trust them to properly use FHE personal data?

> "we literally cannot read your period data."

If the purpose is aggregated data for statistic, then surely the only per-user data they need centrally can already be aggregated (to some degree) on the device, e.g. send back only statistical-distribution variables of the personal data, for distributions over the 3-4 months? And at some point, does the service need to keep collecting data, once the model is good enough (at predicting ovulation etc)?

Another concern would be: If they are building a model, using user data, why should they own the model and thus monetise it (i.e. sell it back to its users) when users get no compensation for supplying that data in the first place.

A flow-tracking app should just stick to that, and purchase the model (for a fee) from a third party. The third party should concern itself with how to get the data without being able to leverage its position as a flow-app maintainer to trick or mislead the majority of its users into giving them free data.

Re: Google is making private AI practical with homomorphic encryption

#262

Great, private AI, at the cost of >1000x the resource usage. Because apparently AI companies weren't already using quite enough energy to cook the planet. The most private AI is the one running on my own hardware, not in some giant data center.

> The most private AI is the one running on my own hardware, not in some giant data center. I want that too, but you gotta ask yourself the question how efficient that is compared to running it in a datacenter shared with everybody else.

Energy efficient, yes, but when you want to keep a query/data private it's maybe worth the extra $ KW. Chicken pie recipes and google AI-search can still go though the datacentres.

As an aside: The computation might also not be the same e.g. ever-changing hidden pre-prompts, security/safety checks blocking or degrading responses, unavoidable verbosity to simple questions, watermarking, collection of prompt data to build user profiles for the purpose of advertising - and we haven't even seen in-response adverts, or sponsor-biased responses yet, but no doubt it's coming.

Re: Google is making private AI practical with homomorphic encryption

#263
post #225

Earlier quoted context omitted.

Then the app can make a bog-standard encrypted-at-rest backup to somewhere and make all the computations on the device on the cleartext data. I don't see the need to do computations on the encrypted data here, which is what FHE would provide in addition to traditional encryption. > and don't have access to the account anymore. This would be trouble with or without FHE. Even if the backend wouldn't need to decrypt the…

Okay, so the platform becomes valuable to its users when it's able to suggest things like "based on millions of users, people with cycles like yours typically ovulate around day 16." In order to do the data mining in order to make those kinds of claims, traditionally you'd need to have access to the data. As you point out, encrypted-at-rest is solved. But what about when it's not at rest? In-use and in-transit is whe…

> Okay, so the platform becomes valuable to its users when it's able to suggest things like "based on millions of users, people with cycles like yours typically ovulate around day 16."

I don't know much about period tracking apps, but is this really the main reason people install those apps? Wouldn't you be able to get similar results by simply monitoring (on-device) the cycle of the person who uses the app for a few months?

How do those apps work before they have millions of users?

All the warnings I've seen about period tracking apps were about unexpected data collection of the entered data. This would be pretty silly if the data collection was integral to what the user expects the app to do.

> Sure, you could just do it locally, but then you miss out on the aggregate data mining.

Ok, a bit of a technical question about FHE here: My understanding of FHE was that you have input data encrypted with some key (plus auxiliary inputs, if needed, that are not encrypted), then you do operations on that data and get a result that is (still) encrypted by that same key.

No questions there as long as you're dealing with a single key.

But the whole point of aggregation and data mining is to combine data from many different users, i.e. inputs that are encrypted by many different keys. Does that work with FHE at all? And if yes, by which key is the aggregation result encrypted?

I don't see how that would work without either "moving" data from one key to another - which would be practically equivalent to decryption - or getting a result that is simultaneously encrypted by all user keys, i.e. practically useless because no one could individually decrypt it.

> The compelling product claim is "we literally cannot read your period data." Not "we pinky swear not to" but "we actually really really actually can't!"

You could obviously read the data enough to do aggregations on it.

If you can do that for "good" purposes, what stops you to use the same aggregation algorithm for advertisers - except pinky promises again?

Re: Google is making private AI practical with homomorphic encryption

#264

My master's thesis is on a topic in this field (Privacy Preserving ML) and from my understanding HE and other techniques have very high overheads(~10^3) on inference tasks and thus aren't very commercially viable.

Do you mind sharing your master's thesis? oO

Would be interesting to read it (and no judgement!)

Re: Google is making private AI practical with homomorphic encryption

#265
It seems to me this tech still presumes the data sits in data warehouses, which i don't like to start with. The tech succeeds in packing the data in identical black boxes, so they all seem equal to the map-reduce function that runs over them. A separate identification layer knows which of the boxes is yours. But who knows, maybe from the results you can be fingerprinted anyway. As in the game, how many questions before you can guess the thing I'm thinking of? Answer: not that many.

Re: Google is making private AI practical with homomorphic encryption

#267
post #194
post #35

Maybe I'm not understanding this, but how is it that you can know enough about the data to process it without undermining the fundamental concept of encryption? Isn't encrypted data supposed to be just random noise without the key? The more you know about the underlying data the easier it gets to decrypt? Does this mean someone can just steal your encrypted data and use that to steal your identity without even needin…

the basic encryption scheme used here is fairly straightforward actually, at least the symmetric encryption version. Let s be a uniformly random, 512-dimensional u32 vector. To encrypt a message m (say a 512-dimensional bit vector for simplicity), you 1. generate a 512 x 512 random (u32) matrix A, and 2. generate a 512-dimensional rounded (to the nearest integer) Gaussian, say of standard deviation 10, e. The ciphert…

Will the new (summed) A, e and b be the same size as the originals, and is m2 + m2 still a 512-dimensional bit vector?

I though (when I tried to understand it) that some part of the HE inflates some component of the result?

Re: Google is making private AI practical with homomorphic encryption

#268
post #47

This is the same Google that doesn't have e2ee on their password manager by default. Like WTF, it's a password manager.

F. Scott Fitzgerald's test of top-tier intelligence - > Holding two opposing views in the mind means accepting two contradictory ideas at the same time without needing to pick one side or rush to a simple answer I continue to use Apple products because they are top class even though everytime I think of Tim Cook in the Oval Office presenting the gold plaque to the current president, it makes me wanna puke. World isnt…

Ah, Google must be very intelligent indeed, then.

Re: Google is making private AI practical with homomorphic encryption

#269
post #203

This is the same Google that doesn't have e2ee on their password manager by default. Like WTF, it's a password manager.

Former Googler here. E2EE is easy. Nobody gets promoted at Google for solving easy problems. In fact if you set out to solve an easy problem, it looks bad at performance review time.

[deleted]

Re: Google is making private AI practical with homomorphic encryption

#270

So much inefficiency just to run it on someone else's untrusted hardware. Private AI is already possible today with local open-weight models running on hardware you control. Homomorphic encryption is cool technology, but I'm really not sure what problem it solves.

I bet this would've been ground breaking if this was an announcement from Apple though.
Post reply on HN