Live data from Hacker News

Why won’t Google give an answer on whether Bard was trained on Gmail data?

skiff.com

41–50 of 74 posts

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#42
post #17
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

Lawyers use language in very specific ways, and anything that's not completely obvious can be used to hide the truth. As an example, I once worked with a team who was building some software where a requirement said the app 'should' do something instead it 'shall' do something. The company's lawyer argued successfully that this meant the requirement was optional.

The Highway Code in the UK is full of “must” and “should” indicating firm requirements and optional choices. The use of “should” is a softer method of persuasion, in the same way signs say “Please close the door” not just “Close the door” which is an instruction. This aligns with the famous British politeness - or at least that is how I interpret the wording.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#43

Earlier quoted context omitted.

I would think that part 4 and "what bard has to say about this" sections would make most people question Google's comment. 4. Google has never denied that Bard was trained on data from Gmail. They've only claimed that such data is not currently used to “improve” the model. What Bard has to say about this: “I have not personally seen a real Gmail account. However, I have access to a massive dataset of Gmail emails, an…

> Bard is an early experiment based on Large Language Models and will make mistakes. It is not trained on Gmail data. -JQ How exactly are you interpreting that statement?

I am defining it on the points I mentioned. I don't know if they are in the wrong or not, but their responses give me pause

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#44
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

JQ is part of the PR team, not engineering. the article author is correct not taking what he says at face value. lots of doublespeak in PR.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#45
post #31

Earlier quoted context omitted.

I would think that part 4 and "what bard has to say about this" sections would make most people question Google's comment. 4. Google has never denied that Bard was trained on data from Gmail. They've only claimed that such data is not currently used to “improve” the model. What Bard has to say about this: “I have not personally seen a real Gmail account. However, I have access to a massive dataset of Gmail emails, an…

They could easily have done this on a subset of consenting users, eg: their own employees.

You are 100% correct. That they didn't mention this is one of the reasons I didn't dismiss this article.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#47
post #38
post #5

Earlier quoted context omitted.

Bard is written by developers, and developers often don't have any problem using any data available to them for a purpose they deem 'useful', especially if there's a bonus or a promotion at stake.

Developers don't make the decisions at Google.

Not anymore, that is.

See also: .zip domain

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#48
post #8

The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

Using the present tense to answer a past-tense question is hardly unambiguous. "Did you send an email to Fred?" -- "No, I'm not sending an email to Fred" doesn't answer the question. It's not a "conspiracy theory" to have realized that big corps have teams of people to frame their public statements with carefully chosen words to present issues in the best light for them, even if it's deeply misleading.

> "Did you send an email to Fred?" -- "No, I'm not sending an email to Fred"

To me this unambiguously says both that I did not send an email to him and I don't intend to. Is it really ambiguous to you?

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#49
post #36
post #24

Earlier quoted context omitted.

They also have added replies which have been removed, reading the article it's actually not as conspiratorial as you make it seem. In this case I really think it prudent to assume the worst from Google as they don't really have a positive history for walling off users data, be it personal email or phone meta information. > "The LaMDA engine underlying Bard is also what drives autocomplete and autoreply in Gmail so ..…

Who wrote that reply? A Googler? Is there a screenshot? This would be a gigantic GDPR lawsuit. This reminds me of not communicating with Gmail users. Gmail has been evil forever since "personalized" ads.

It's in the article, but here's the direct link to help you out.

https://twitter.com/cajundiscordian/status/16382433030356705...

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#50
post #36
post #24

Earlier quoted context omitted.

They also have added replies which have been removed, reading the article it's actually not as conspiratorial as you make it seem. In this case I really think it prudent to assume the worst from Google as they don't really have a positive history for walling off users data, be it personal email or phone meta information. > "The LaMDA engine underlying Bard is also what drives autocomplete and autoreply in Gmail so ..…

Who wrote that reply? A Googler? Is there a screenshot? This would be a gigantic GDPR lawsuit. This reminds me of not communicating with Gmail users. Gmail has been evil forever since "personalized" ads.

At the moment, for some reason Google seems almost untouchable.

I mean, single handedly destroying the browser market by deceit and abuse of market position in broad daylight, you'd think sooner or later EU or someone would force them to pay and put up a browser ballot on Google.com, but so far, no.

Luckily the French consumer protection agency has at least forced them to implement the cookie question thing almost so we can now reject all right away.

Post reply on HN