Live data from Hacker News

Why won’t Google give an answer on whether Bard was trained on Gmail data?

skiff.com

21–30 of 74 posts

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#21
post #17
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

Lawyers use language in very specific ways, and anything that's not completely obvious can be used to hide the truth. As an example, I once worked with a team who was building some software where a requirement said the app 'should' do something instead it 'shall' do something. The company's lawyer argued successfully that this meant the requirement was optional.

I'm not sure what you're trying to say, but "It is not trained on Gmail data." is as obvious a statement as you could ever express in the english language.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#22
post #17
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

Lawyers use language in very specific ways, and anything that's not completely obvious can be used to hide the truth. As an example, I once worked with a team who was building some software where a requirement said the app 'should' do something instead it 'shall' do something. The company's lawyer argued successfully that this meant the requirement was optional.

Not only that but this tweet is from Google Workspace and so it is not at all unreasonable to say it is not trained on Google Workspace Gmail and it says nothing about public gmail. Words. We have them.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#23
post #17

Earlier quoted context omitted.

Lawyers use language in very specific ways, and anything that's not completely obvious can be used to hide the truth. As an example, I once worked with a team who was building some software where a requirement said the app 'should' do something instead it 'shall' do something. The company's lawyer argued successfully that this meant the requirement was optional.

That's how Internet standards work. https://www.ietf.org/rfc/rfc2119.txt

This was a contract between two companies, written by a product manager and a CEO. It wasn't a technical RFC.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#24
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

They also have added replies which have been removed, reading the article it's actually not as conspiratorial as you make it seem.

In this case I really think it prudent to assume the worst from Google as they don't really have a positive history for walling off users data, be it personal email or phone meta information.

> "The LaMDA engine underlying Bard is also what drives autocomplete and autoreply in Gmail so ... yeah Bard's training data includes Gmail. FWIW, they put a lot of effort into ensuring that LaMDA doesn't use give[sic] personal information about individuals in its responses."

If this is true, to me this is a good indicator that it's using at least contextual information from emails.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#26
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

Using the present tense to answer a past-tense question is hardly unambiguous.

"Did you send an email to Fred?" -- "No, I'm not sending an email to Fred" doesn't answer the question.

It's not a "conspiracy theory" to have realized that big corps have teams of people to frame their public statements with carefully chosen words to present issues in the best light for them, even if it's deeply misleading.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#28
post #4

The article directly quotes google giving a clear answer…. “Google replied to the tweet directly, saying, “Bard is an early experiment based on Large Language Models and will make mistakes. It is not trained on Gmail data. -JQ”.

Elsewhere in this thread we noted how peculiar legal language can be.

1. "is" versus "was" 2. This tweet is from Google Workspace. It might pertain to Google Workspace gmail only and say nothing about public gmail.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#29
post #23

Earlier quoted context omitted.

That's how Internet standards work. https://www.ietf.org/rfc/rfc2119.txt

This was a contract between two companies, written by a product manager and a CEO. It wasn't a technical RFC.

This choice of wording in systems engineering is not ambiguous, and should effectively means optional. These words are often in capital letters trying to highlight the importance of it. It is absolutely not limited to RFCs, and is often used in a software specification.

Product manager and CEO _should_ know better. It's very understandable that they don't - and I have empathy for them, but unfortunately they're wrong.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#30
Why, or how, would bars know what it is trained on?

And even if bard contains training data recent enough to include discussions on bard, how would it be able to tell speculation from facts?

The only way I can think of is through deliberate alignment. After training on source data, the model is fine tuned by human curated chat dialogue.

Post reply on HN