Live data from Hacker News

I caught Google Gemini using my data and then covering it up

unbuffered.stream

81–87 of 87 posts

Re: I caught Google Gemini using my data and then covering it up

#81
post #5

>But why is Gemini instructed not to divulge its existence? Seems like a reasonable thing to add. Imagine how impersonal chats would feel if Gemini responded to "what food should I get for my dog?" with "according to your `user_context`, you have a husky, and the best food for him is...". They're also not exactly hiding the fact that memory/"personalization" exists either: https://blog.google/products/gemini/temporar…

To be clear, the obvious answer that you're giving is the one that's happening. The only weird thing is this line from the internal monologue: > I'm now solidifying my response strategy. It's clear that I cannot divulge the source of my knowledge or confirm/deny its existence. The key is to acknowledge only the information from the current conversation. Why does it think that it's not allowed to confirm/deny the exis…

its not allowed to confirm/deny security; privacy; copyright and IP violations.

Re: I caught Google Gemini using my data and then covering it up

#82
I saw something like this in ChatGPT in the spring when it refused to tell me something about a keyboard emulator (USB Rubber Ducky) because it was unethical, but then looking at the thinking gave me the answer.

Shocked you can still exploit this. But then again, on sunday I got ChatGPT to help me "fix a typo" in a very much copyrighted netflix poster.

Re: I caught Google Gemini using my data and then covering it up

#83
post #81

Earlier quoted context omitted.

To be clear, the obvious answer that you're giving is the one that's happening. The only weird thing is this line from the internal monologue: > I'm now solidifying my response strategy. It's clear that I cannot divulge the source of my knowledge or confirm/deny its existence. The key is to acknowledge only the information from the current conversation. Why does it think that it's not allowed to confirm/deny the exis…

its not allowed to confirm/deny security; privacy; copyright and IP violations.

Aside: I'm well aware that it's just about as popular to use an Oxford comma than to not, but this might be the first time I've ever seen someone omit an Oxford semicolon as it really seems odd to me. But Automatic Semicolon Insertion in Javascript is odd to me, as well.

Re: I caught Google Gemini using my data and then covering it up

#84
post #81

Earlier quoted context omitted.

its not allowed to confirm/deny security; privacy; copyright and IP violations.

Aside: I'm well aware that it's just about as popular to use an Oxford comma than to not, but this might be the first time I've ever seen someone omit an Oxford semicolon as it really seems odd to me. But Automatic Semicolon Insertion in Javascript is odd to me, as well.

To be fair, the Oxford semicolon might be incorrect here if they intended "copyright and IP violations" to be paired.

e.g. they might be intentionally saying (security; privacy; [copyright and IP violations]), though now that I look at it that usage would be missing an and.

Re: I caught Google Gemini using my data and then covering it up

#85
post #81

Earlier quoted context omitted.

its not allowed to confirm/deny security; privacy; copyright and IP violations.

Aside: I'm well aware that it's just about as popular to use an Oxford comma than to not, but this might be the first time I've ever seen someone omit an Oxford semicolon as it really seems odd to me. But Automatic Semicolon Insertion in Javascript is odd to me, as well.

now you know im human ;)

Re: I caught Google Gemini using my data and then covering it up

#86

Earlier quoted context omitted.

Aside: I'm well aware that it's just about as popular to use an Oxford comma than to not, but this might be the first time I've ever seen someone omit an Oxford semicolon as it really seems odd to me. But Automatic Semicolon Insertion in Javascript is odd to me, as well.

To be fair, the Oxford semicolon might be incorrect here if they intended "copyright and IP violations" to be paired. e.g. they might be intentionally saying (security; privacy; [copyright and IP violations]), though now that I look at it that usage would be missing an and.

I'm fine with an implied "and," but I disagree that "violations" is scoped only to the final item in the series. It's clearly transitive to each item in the series (security violations, etc.). That said, your point stands that the phrase "copyright and IP" could be the final item in the series (with an omitted "and" at the series level) rather than the final two items, although there wouldn't be a compelling reason to do that in this particular case.

Re: I caught Google Gemini using my data and then covering it up

#87
post #5

>But why is Gemini instructed not to divulge its existence? Seems like a reasonable thing to add. Imagine how impersonal chats would feel if Gemini responded to "what food should I get for my dog?" with "according to your `user_context`, you have a husky, and the best food for him is...". They're also not exactly hiding the fact that memory/"personalization" exists either: https://blog.google/products/gemini/temporar…

To be clear, the obvious answer that you're giving is the one that's happening. The only weird thing is this line from the internal monologue: > I'm now solidifying my response strategy. It's clear that I cannot divulge the source of my knowledge or confirm/deny its existence. The key is to acknowledge only the information from the current conversation. Why does it think that it's not allowed to confirm/deny the exis…

Think about it. The chatbot has found itself in a scenario where it appears to be acting maliciously. This isn't actually true, but the user's response has made it seem this way. This lead it to completely misunderstand the intention of the instruction in the system prompt.

So what is the natural way for this scenario to continue? To inexplicably come clean, or to continue acting maliciously? I wouldn't be surprised if in such a scenario it started acting malicious in other unrelated ways just because that is what it thinks is a likely way for the conversation to continue

Post reply on HN