Live data from Hacker News

Why won’t Google give an answer on whether Bard was trained on Gmail data?

skiff.com

31–40 of 74 posts

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#31
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

I would think that part 4 and "what bard has to say about this" sections would make most people question Google's comment. 4. Google has never denied that Bard was trained on data from Gmail. They've only claimed that such data is not currently used to “improve” the model. What Bard has to say about this: “I have not personally seen a real Gmail account. However, I have access to a massive dataset of Gmail emails, an…

They could easily have done this on a subset of consenting users, eg: their own employees.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#32
It's a bad idea for a company to comment on a specific case even if they've done nothing wrong, because it creates the expectation that the company will comment when they are in the clear. And complicit otherwise.

It's safer for companies to not comment at all, unless forced to.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#33
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

I would think that part 4 and "what bard has to say about this" sections would make most people question Google's comment. 4. Google has never denied that Bard was trained on data from Gmail. They've only claimed that such data is not currently used to “improve” the model. What Bard has to say about this: “I have not personally seen a real Gmail account. However, I have access to a massive dataset of Gmail emails, an…

> Bard is an early experiment based on Large Language Models and will make mistakes. It is not trained on Gmail data. -JQ

How exactly are you interpreting that statement?

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#34
post #23

Earlier quoted context omitted.

This was a contract between two companies, written by a product manager and a CEO. It wasn't a technical RFC.

This choice of wording in systems engineering is not ambiguous, and should effectively means optional. These words are often in capital letters trying to highlight the importance of it. It is absolutely not limited to RFCs, and is often used in a software specification. Product manager and CEO _should_ know better. It's very understandable that they don't - and I have empathy for them, but unfortunately they're wrong…

_must_ tbqh lmao

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#35
post #23

Earlier quoted context omitted.

This was a contract between two companies, written by a product manager and a CEO. It wasn't a technical RFC.

This choice of wording in systems engineering is not ambiguous, and should effectively means optional. These words are often in capital letters trying to highlight the importance of it. It is absolutely not limited to RFCs, and is often used in a software specification. Product manager and CEO _should_ know better. It's very understandable that they don't - and I have empathy for them, but unfortunately they're wrong…

Sure, and hence the company won. The point here is that language can have very specific meaning. There is a difference between 'should' and 'shall', even though most people would think they're effectively the same. If I said "You should complete that task" to one of my team I'm not really giving them the choice and leaving it up to them. I'm telling them to do something. Outside of RFCs and contracts language is a bit ambiguous and lawyers use that to their advantage.

In exactly the same way, there is also a difference between 'is' and 'was', and I think it's totally plausible that a Google lawyer might use that to hide the fact they used GMail data to train AI in the past.

That doesn't mean they did. It only means I wouldn't be surprised if someone proves they did, and that their lawyer used the tense of a response to try to hide it.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#36
post #24
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

They also have added replies which have been removed, reading the article it's actually not as conspiratorial as you make it seem. In this case I really think it prudent to assume the worst from Google as they don't really have a positive history for walling off users data, be it personal email or phone meta information. > "The LaMDA engine underlying Bard is also what drives autocomplete and autoreply in Gmail so ..…

Who wrote that reply? A Googler? Is there a screenshot? This would be a gigantic GDPR lawsuit.

This reminds me of not communicating with Gmail users. Gmail has been evil forever since "personalized" ads.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#37
post #6

Would this not be fairly easy to test? Find some fairly dense thing in gmail, stick 50% into bard and ask it to complete rest& see how close output is?

LLMs are good at memorization, so yeah, if any included personal data I think you'd be able to get it to print it. (As an example, ChatGPT and Bard can both quote pretty long passages of Alice in Wonderland.) There aren't any techniques I know of to prevent it either; when training an image model the recommendation is to dedupe the input so nothing is weighted over anything else, but that's not an absolute defense.

Alice is many times in the dataset.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#38
post #5

Logically it just doesn't have that data and why would it? If you had proprietary data you didn't want to share why would you add it to the language model?

Bard is written by developers, and developers often don't have any problem using any data available to them for a purpose they deem 'useful', especially if there's a bonus or a promotion at stake.

Developers don't make the decisions at Google.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#39
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

When a corporation answers in a specific way it's because it has a specific meaning. 'Ooops we meant 'x' doesn't tend to hold up so well in court.' Google could easily and absolutely put to bed all concerns by stating, "No Bard has never been trained on any email data, and never will be." Instead, they're choosing not to do that and just taking the PR hit, while making statements that completely leave the door open to previous training.
Post reply on HN