Live data from Hacker News

Employees are feeding sensitive data to ChatGPT, raising security fears

darkreading.com

71–80 of 355 posts

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#71

ChatGPT Business Edition seems pretty obvious and I'd surprised if OpenAI isn't already working on it. Separate models for each customer, data silos and protection. The infra is already there on Azure.

It actually is on Azure, exactly as you described.

https://learn.microsoft.com/en-us/azure/cognitive-services/o...

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#72
post #5

We block ChatGPT, as do most federal contractors. I think it’s a horrible exploit waiting to happen: - there’s no way they’re manually scrubbing out sensitive data so its bound to spill out from the training data when prompting the model - OpenAI is openly storing all this data they’re collecting to the extent that they’ve had several leaks now where people can see others’ conversations and data. We are one step away…

If your competitor use ChatGPT to compete with you and they're 10x productive than yours, are you still willing to insist? If the productive is 100x, will you?

This isn't an argument of ChatGPT vs nothing. This is an argument of "external" ChatGPT vs some other AI sitting on your own secured hardware, maybe even a branch of ChatGPT.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#73

Earlier quoted context omitted.

If your competitor use ChatGPT to compete with you and they're 10x productive than yours, are you still willing to insist? If the productive is 100x, will you?

It might be just as likely that ChapGPT will cause a mistake like Knight Capital because no one bothered to thoroughly verify the AI's looks-good-but-deeply-flawed answer, and the two aren't mutually exclusive possibilities.

Right. I've had ChatGPT completely fail at something as simple as writing a batch file to find and replace text in a text file.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#74

This is nothing new at all. How many people have Grammarly plugins installed? They are advertising aggressively, so I'd think it is the new hotness. Don't tell me Grammarly is not hoovering up all of the Slack, Word, Docs, and Gmail data that everyone sends it, and holding on for some future purpose. We'll see.

Grammarly is an OpSec nightmare that’s somehow managed to slip under most people’s radar. I know folks who won’t use the Okta browser extension because of the extensive permissions it asks for, but will happily use Grammarly on everything.

Last time I looked (maybe things have improved…?) Grammarly would automatically attach itself to any text box you interacted with, and immediately send all of your content to their servers for processing. How this software gets past IT departments is a mystery to me.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#75
post #58

Earlier quoted context omitted.

> Google takes your data and sells it. Literally making your data available to the highest bidder. Even if they are not doing it now(?), what makes you think that they will not do so in the future? It's not like your data has an expiration date.

Because they are completely different business models. If OpenAI decides to become an advertising behemoth then I would show concern. Right now they use your data for training (when they use it).

OpenAI has already demonstrated that they're all in for maximizing profit. They may not be advertisers, but advertisers aren't the only sorts of companies that make bank by selling personal data.

I see no reason to think OpenAI would leave that money on the table.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#76
This is one of the reasons Databricks created Dolly, a slim LLM that unlocks the magic of ChatGPT. A homegrown LLM that can tap into/query the datasets of all the data in an organizations Data Lakehouse will be hugely powerful.

I am working with customers that are looking to train a homegrown LLM that they host and have blocked access to ChatGPT.

https://www.datanami.com/2023/03/24/databricks-bucks-the-her...

https://news.ycombinator.com/item?id=35288063

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#77
post #42

We went pretty quickly from: No way I’m giving Google any of my data! I will use 5 different browsers in incognito mode and never log in. To -> Sure I will login with my name and email and feed you as much of my most personal thoughts and data as I can dear ChatGPT!

Google takes your data and sells it. Literally making your data available to the highest bidder. Is OpenAI doing that? If Google existed in its current form during the early internet it would be classified in the same category as Bonzai Buddy. Spyware. That is what Google is. So I can very reasonably understand why people would trust OpenAI with data they wouldn't trust Google with. OpenAI hasn't spit in the face of…

> Google takes your data and sells it. Literally making your data available to the highest bidder.

But it doesn't, does it? It sells the fact that it knows everything about everyone and can get any ad to the perfect people for it. It's not going on the open market and telling people I regularly buy 12 lbs of marshmallow fluff and then use it in videos I keep on my google drive.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#78
We saw these same fears with the release of Gmail. Why would you trust your email to Google?!! Aren't they going to train their spam filters on all your data? Aren't they going to sell it, or use it to sell you ads?

Corporations constantly put their most sensitive data in 3rd party tools. The executive in the article was probably copying his company strategy from Google docs.

Yes, there are good reasons for concern, but the power of the tool is simply too great to ignore.

Banning these tools will go the same way as prohibition did in the US, people will simply ignore it until it becomes too absurd to maintain and too profitable to not participate in.

Companies which are able to operate without these fears will move faster, grow more quickly, and ultimately challenge companies restricted to operate without.

Now I think the article should be a wake-up call for OpenAI. Messaging around what is and what is not used for training could be improved. Corporate accounts for Chat with clearer privacy policies would be great and warnings that, yes, LLMs do memorize data and you should treat anything you put into a free product on the web as fair game for someone's training algorithm.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#79
post #5

We block ChatGPT, as do most federal contractors. I think it’s a horrible exploit waiting to happen: - there’s no way they’re manually scrubbing out sensitive data so its bound to spill out from the training data when prompting the model - OpenAI is openly storing all this data they’re collecting to the extent that they’ve had several leaks now where people can see others’ conversations and data. We are one step away…

So you block internet access for all employees? Cos anything you think is being pasted into ChatGPT is being pasted everywhere, whether its Google, Slack, Chrome Plugins, public Wifi.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#80
post #5

We block ChatGPT, as do most federal contractors. I think it’s a horrible exploit waiting to happen: - there’s no way they’re manually scrubbing out sensitive data so its bound to spill out from the training data when prompting the model - OpenAI is openly storing all this data they’re collecting to the extent that they’ve had several leaks now where people can see others’ conversations and data. We are one step away…

Possibly I don't know how this all works, but I think if the host of a ChatGPT interface were willing to provide their own API key (and pay), they could then provide a "service" to others (and collect all input). In that case, you wouldn't know to block them until it was too late. Ultimately either you must watch/block all outgoing traffic, or you must train your people so thoroughly that they become suspicious of ev…

> Possibly I don’t know how this all works, but I think if the host of a ChatGPT interface were willing to provide their own API key (and pay), they could then provide a “service” to others (and collect all input).

Well, GP was referring to blocking ChatGPT as a federal contractor. I suspect that as a federal contractor, they are also vetting other people that they share data with, not just blocking ChatGPT as a one-off thing. I mean, generic federal data isn’t as tightly regulated as, say, HIPAA PHI (having spent quite a lot of time working for a place that handles both), but there are externally-imposed rules and consequences, unlike simple internal-proprietary data.

Post reply on HN