The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").
I would think that part 4 and "what bard has to say about this" sections would make most people question Google's comment. 4. Google has never denied that Bard was trained on data from Gmail. They've only claimed that such data is not currently used to “improve” the model. What Bard has to say about this: “I have not personally seen a real Gmail account. However, I have access to a massive dataset of Gmail emails, an…
Why won’t Google give an answer on whether Bard was trained on Gmail data?
31–40 of 74 posts
Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?
#32It's safer for companies to not comment at all, unless forced to.
Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?
#33The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").
I would think that part 4 and "what bard has to say about this" sections would make most people question Google's comment. 4. Google has never denied that Bard was trained on data from Gmail. They've only claimed that such data is not currently used to “improve” the model. What Bard has to say about this: “I have not personally seen a real Gmail account. However, I have access to a massive dataset of Gmail emails, an…
How exactly are you interpreting that statement?
Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?
#34Earlier quoted context omitted.
This was a contract between two companies, written by a product manager and a CEO. It wasn't a technical RFC.
This choice of wording in systems engineering is not ambiguous, and should effectively means optional. These words are often in capital letters trying to highlight the importance of it. It is absolutely not limited to RFCs, and is often used in a software specification. Product manager and CEO _should_ know better. It's very understandable that they don't - and I have empathy for them, but unfortunately they're wrong…
Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?
#35Earlier quoted context omitted.
This was a contract between two companies, written by a product manager and a CEO. It wasn't a technical RFC.
This choice of wording in systems engineering is not ambiguous, and should effectively means optional. These words are often in capital letters trying to highlight the importance of it. It is absolutely not limited to RFCs, and is often used in a software specification. Product manager and CEO _should_ know better. It's very understandable that they don't - and I have empathy for them, but unfortunately they're wrong…
In exactly the same way, there is also a difference between 'is' and 'was', and I think it's totally plausible that a Google lawyer might use that to hide the fact they used GMail data to train AI in the past.
That doesn't mean they did. It only means I wouldn't be surprised if someone proves they did, and that their lawyer used the tense of a response to try to hide it.
Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?
#36The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").
They also have added replies which have been removed, reading the article it's actually not as conspiratorial as you make it seem. In this case I really think it prudent to assume the worst from Google as they don't really have a positive history for walling off users data, be it personal email or phone meta information. > "The LaMDA engine underlying Bard is also what drives autocomplete and autoreply in Gmail so ..…
This reminds me of not communicating with Gmail users. Gmail has been evil forever since "personalized" ads.
Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?
#37Would this not be fairly easy to test? Find some fairly dense thing in gmail, stick 50% into bard and ask it to complete rest& see how close output is?
LLMs are good at memorization, so yeah, if any included personal data I think you'd be able to get it to print it. (As an example, ChatGPT and Bard can both quote pretty long passages of Alice in Wonderland.) There aren't any techniques I know of to prevent it either; when training an image model the recommendation is to dedupe the input so nothing is weighted over anything else, but that's not an absolute defense.
Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?
#38Logically it just doesn't have that data and why would it? If you had proprietary data you didn't want to share why would you add it to the language model?
Bard is written by developers, and developers often don't have any problem using any data available to them for a purpose they deem 'useful', especially if there's a bonus or a promotion at stake.
Re: Why won’t Google give an answer on whether Bard was trained on Gmail data?
#39The article decides to mention that Google made a public statement that clearly and unambiguously answers their question ( https://twitter.com/GoogleWorkspace/status/16382985371956019... ) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").