Live data from Hacker News

I caught Google Gemini using my data and then covering it up

unbuffered.stream

21–30 of 87 posts

Re: I caught Google Gemini using my data and then covering it up

#21
post #5

>But why is Gemini instructed not to divulge its existence? Seems like a reasonable thing to add. Imagine how impersonal chats would feel if Gemini responded to "what food should I get for my dog?" with "according to your `user_context`, you have a husky, and the best food for him is...". They're also not exactly hiding the fact that memory/"personalization" exists either: https://blog.google/products/gemini/temporar…

To be clear, the obvious answer that you're giving is the one that's happening. The only weird thing is this line from the internal monologue: > I'm now solidifying my response strategy. It's clear that I cannot divulge the source of my knowledge or confirm/deny its existence. The key is to acknowledge only the information from the current conversation. Why does it think that it's not allowed to confirm/deny the exis…

Could be that it’s confusing not mentioning the literal term “user_context” vs the existence of it. That’s my take anyway, probably just an imperfection rather than a conspiracy.

Re: I caught Google Gemini using my data and then covering it up

#22

This is a fundamental violation of trust. If an AI llm is meant to eventually evolve into general intelligence capable of true reasoning, then we are essentially watching a child grow up. Posts like this are screaming "you're raising a psychopath!!"... If AI is just an overly complicated a stack of autocorrect functions, this proves its behavior heavily if not entirely swayed by its usually hidden rules to the point…

LLM's are not kids. Kids sometimes lie, it's a part of the learning process. Lying to cover up a mistake is not a strong sign of psychopathy.

> This is a fundamental violation of trust.

I don't disagree. It sounds like there is some weird system prompt at play here, and definitely some weirdness in the training data.

Re: I caught Google Gemini using my data and then covering it up

#25
post #15

Okay, this is a weird place to "publish" this information, but I'm feeling lazy, and this is the most of an "audience" I'll probably have. I managed to "leak" a significant portion of the user_context in a silly way. I won't reveal how, though you can probably guess based on the snippets. It begins with the raw text of recent conversations: > Description: A collection of isolated, raw user turns from past, unrelated…

Oh is this the famous "I got Google ads based on conversations it must have picked up from my microphone"?

Re: I caught Google Gemini using my data and then covering it up

#26
post #5

>But why is Gemini instructed not to divulge its existence? Seems like a reasonable thing to add. Imagine how impersonal chats would feel if Gemini responded to "what food should I get for my dog?" with "according to your `user_context`, you have a husky, and the best food for him is...". They're also not exactly hiding the fact that memory/"personalization" exists either: https://blog.google/products/gemini/temporar…

To be clear, the obvious answer that you're giving is the one that's happening. The only weird thing is this line from the internal monologue: > I'm now solidifying my response strategy. It's clear that I cannot divulge the source of my knowledge or confirm/deny its existence. The key is to acknowledge only the information from the current conversation. Why does it think that it's not allowed to confirm/deny the exis…

One explanation might be if the instruction was "under no circumstances mention user_context unless the user brings it up" and technically the user didn't bring it up, they just asked about the previous response.

Re: I caught Google Gemini using my data and then covering it up

#27
These things aren't conspiracies. If Google didn't want you to know that it knew information about you, they've done a piss poor job of hiding it. Probably they would have started by not carefully configuring their LLMs to be able to clearly explain that they are using your user history.

Instead, the right conclusion is: the LLM did a bad job with this answer. LLMs often provide bad answers! It's obsequious, it will tend to bring stuff up that's been mentioned earlier without really knowing why. It will get confused and misexplain things. LLMs are often badly wrong in ways that sound plausibly correct. This is a known problem.

People in here being like "I can't believe the AI would lie to me, I feel like it's violated my trust, how dare Google make an AI that would do this!" It's an AI. Their #1 flaw is being confidently wrong. Should Google be using them here? No, probably not, because of this fact! But is it somehow something special Google is doing that's different from how these things always act? Nope.

Re: I caught Google Gemini using my data and then covering it up

#28
post #14

I’m pretty sure this is because they don’t want Gemini saying things like, “based on my stored context from our previous chat, you said you were highly proficient in Alembic.” It’s hard to get a principled autocomplete system like these to behave consistently. Take a look at Claude’s latest memory-system prompt for how it handles user memory. https://x.com/kumabwari/status/1986588697245196348

Yeah but what if you explicitly ask it, "what/how do you know about my stored context"? Why should it be instructed to lie then?

It could be that the instruction was vague enough ("never mention user_context unless the user brings it up", eg) and since the user never mentioned "context", the model treated it as not having been, technically speaking, mentioned.
Post reply on HN