Live data from Hacker News

Google's Gemini AI caught scanning Google Drive PDF files without permission

tomshardware.com

61–70 of 163 posts

Re: Google's Gemini AI caught scanning Google Drive PDF files without permission

#61
There is a fundamentally interesting nuance to highlight. I don't know precisely what google is doing, but if they're just shuttling the content through a closed-loop deterministic LLM, then, much like a spellchecker, I see no issue. Sure, it _feels_ creepy, but it's just an algo.

Perhaps someone can articulate the precise threshold of 'access' they wish to deny apps that we overtly use? And how would that threshold be defined?

"Do not run my content through anything more complicated than some arbitrary [complexity metric]" ??

Re: Google's Gemini AI caught scanning Google Drive PDF files without permission

#62

Earlier quoted context omitted.

Just because “the cloud” is someone else’s computer doesn’t mean it doesn’t exist.

I think that wasn’t supposed to be taken literally but more tongue in cheek. The main point being that it belongs to some other party. But the cloud buzzword is fuzzy in description. Ever since ‘cloud’ privacy took a nosedive

It's worth noting that cloud computing has existed since the 1960s. It just used to be called "time-sharing".

Re: Google's Gemini AI caught scanning Google Drive PDF files without permission

#63
post #52
post #30

Earlier quoted context omitted.

Models can easily regurgitate back training data verbatim, so anything private can be in theory accessed by anyone without proper access to that file

This is partly true but less and less every day. IMO the bigger concern is that this data is not just used to train models. It is stored, completely verbatim, in the training set data. They aren’t pulling from PDFs in realtime during training runs, they’re aggregating all of that text and storing it somewhere. And that somewhere is prone to employees viewing, leakage to the internet, etc.

> This is partly true but less and less every day.

Isn't this like encryption, though?

I'm fairly sure that the cryptography community basically says: if someone has a copy of your encrypted data for a long time, the likelihood over time for them to be able to read it approaches 100%, regardless of the current security standard you're using.

Who could possibly guarantee that whatever LLM is safe now will be safe at all times over the next 5-10-20 years? And if they're guaranteeing, they're lying.

Re: Google's Gemini AI caught scanning Google Drive PDF files without permission

#64
post #28

Every single week I have to refuse enabling back up for my pictures on my Google pixel. I refuse it today, next week I open the app and the UI shows the back up option enabled with a button "continue using the app with back up". Somebody took the time to talk down my comment about this being a strategy to give their AI more training data. I continue believing that if they have your data they will use it.

I in no way want to absolve Google, but that's the case for so many app permissions on Android. Turn off notifications, and two weeks later the same app your turned off notifications for is once again sending you notifications. It's beyond a joke.

Name sure you also disable the ability for the app to change settings

Re: Google's Gemini AI caught scanning Google Drive PDF files without permission

#65
post #28

Every single week I have to refuse enabling back up for my pictures on my Google pixel. I refuse it today, next week I open the app and the UI shows the back up option enabled with a button "continue using the app with back up". Somebody took the time to talk down my comment about this being a strategy to give their AI more training data. I continue believing that if they have your data they will use it.

I in no way want to absolve Google, but that's the case for so many app permissions on Android. Turn off notifications, and two weeks later the same app your turned off notifications for is once again sending you notifications. It's beyond a joke.

Can you share some apps where this happens for you. I have rather the complete opposite experience where unused apps with permissions eventually lose said permissions.

Re: Google's Gemini AI caught scanning Google Drive PDF files without permission

#67

There is a fundamentally interesting nuance to highlight. I don't know precisely what google is doing, but if they're just shuttling the content through a closed-loop deterministic LLM, then, much like a spellchecker, I see no issue. Sure, it _feels_ creepy, but it's just an algo. Perhaps someone can articulate the precise threshold of 'access' they wish to deny apps that we overtly use? And how would that threshold…

It was already possible to search for photos in Google Drive by their content. They seemed to be doing some sort of image tagging and feeding that into search results. Did that ever cause a fuss?

I think the more interesting point is how little people seem to care for the auto-summarization feature. Like, why would anyone want to see their archived tax docs summarized by a chatbot? I think whether an "AI" did that or not is almost a red herring.

Re: Google's Gemini AI caught scanning Google Drive PDF files without permission

#68

Meta commentary but still relevant I think: The author first refers to his source as Kevin Bankston in the article's subtitle. This is also the name shown in the embedded tweet. But the following two references call him Kevin _Bankster_ (which seems like an amusing portmanteau of banker and gangster I guess). Is the author not proofreading his own copy? Are there no editors? If the author can't even keep the name of…

There are no editors.

Maybe an AI editor?

That would be somewhat disconcerting.

Write about problem with AI and article get changed to 10 best fried chicken recipes.

Re: Google's Gemini AI caught scanning Google Drive PDF files without permission

#69
this is similar to the scramble for health data during covid where a number of groups tried (and some succeeded) at using the crisis to squeeze the toothpaste out of the tube in a similar way, as there are low costs to being reprimanded and high value in grabbing the data. bureaucratic smash-and-grabs, essentially. disappointing, but predictable to anyone who has worked in privacy, and most people just make a show of acting surprised then moving on because their careers depend on their ability to sustain a gallopingly absurd best-intentions narrative.

your hacked SMS messages from AT&T are probably next, and everyone will be just as surprised when keystrokes from your phones get hit, or there is a collection agent for model training (privacy enhanced for your pleasure, surely) added as an OS update to commercial platforms.

Make an example of the product managers and engineers behind this, or see it done worse and at a larger scale next time.

Re: Google's Gemini AI caught scanning Google Drive PDF files without permission

#70

There is a fundamentally interesting nuance to highlight. I don't know precisely what google is doing, but if they're just shuttling the content through a closed-loop deterministic LLM, then, much like a spellchecker, I see no issue. Sure, it _feels_ creepy, but it's just an algo. Perhaps someone can articulate the precise threshold of 'access' they wish to deny apps that we overtly use? And how would that threshold…

It was already possible to search for photos in Google Drive by their content. They seemed to be doing some sort of image tagging and feeding that into search results. Did that ever cause a fuss? I think the more interesting point is how little people seem to care for the auto-summarization feature. Like, why would anyone want to see their archived tax docs summarized by a chatbot? I think whether an "AI" did that or…

Right, but it's triggered by the user themselves, per the article:

> "[it] only happens after pressing the Gemini button on at least one document"

I agree the AI aspect is largely a red herring. But I don't think running an algo like a spellchecker within an open document is so awful. If people hate it or it's not useful or accurate, then it should be binned, ofc. And if we're ignoring the AI aspect, then it's just a meh/crappy feature. Not especially newsworthy IMHO.

Post reply on HN