Live data from Hacker News

Why won’t Google give an answer on whether Bard was trained on Gmail data?

skiff.com

1–10 of 74 posts

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#5

Logically it just doesn't have that data and why would it? If you had proprietary data you didn't want to share why would you add it to the language model?

Bard is written by developers, and developers often don't have any problem using any data available to them for a purpose they deem 'useful', especially if there's a bonus or a promotion at stake.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#7

Logically it just doesn't have that data and why would it? If you had proprietary data you didn't want to share why would you add it to the language model?

Agree, I doubt it would be trained.

If something personal - like Gmail/ drive goes out in any of bard responses, it will be the end of bard.

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#8
The article decides to mention that Google made a public statement that clearly and unambiguously answers their question (https://twitter.com/GoogleWorkspace/status/16382985371956019...) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#9
post #5

Logically it just doesn't have that data and why would it? If you had proprietary data you didn't want to share why would you add it to the language model?

Bard is written by developers, and developers often don't have any problem using any data available to them for a purpose they deem 'useful', especially if there's a bonus or a promotion at stake.

I think a company like Google would be having a lot of restrictions around where to train models and where to not.

And they already have access to the whole internet, Gmail conversations would be one tiny part of it.

Also wonder if they got to actually train on github data (considering the Microsoft angle)

Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?

#10
post #4

The article directly quotes google giving a clear answer…. “Google replied to the tweet directly, saying, “Bard is an early experiment based on Large Language Models and will make mistakes. It is not trained on Gmail data. -JQ”.

It does, but there’s more discussion beyond that.
Post reply on HN