Live data from Hacker News

Employees are feeding sensitive data to ChatGPT, raising security fears

darkreading.com

261–270 of 355 posts

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#261
For fun I once just made a blank from with a submission button and a giant text field.

It was quite amazing what people would submit unprompted, so I'm not at all surprised that people would feed sensitive data into ChatGPT. The next cycle will be that ChatGPT gets - surprise - trained on that data, and may start using fragments of it - which may well still be sensitive enough to cause trouble - as its output.

Don't paste confidential information into a textbox, in fact don't trust anybody or any company with your confidential information unless there is a strong contractual relationship backed up by penalties if it gets broken. And even then: the ultimate responsibility is yours, you may be able to recover some $ for damages but your reputation may well be toast.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#262

We saw these same fears with the release of Gmail. Why would you trust your email to Google?!! Aren't they going to train their spam filters on all your data? Aren't they going to sell it, or use it to sell you ads? Corporations constantly put their most sensitive data in 3rd party tools. The executive in the article was probably copying his company strategy from Google docs. Yes, there are good reasons for concern,…

I think this is different in that ChatGPT is expressly using your data as training in a probabilistic model. This means: * Their contractors can (and do!) see your chat data to tune the model * If the model is trained on your confidential data, it may start returning this data to other users (as we've seen with Github Copilot regurgitating licensed software) * The site even _tells you_ not to put confidential data in…

> I think this is different in that ChatGPT is expressly using your data as training in a probabilistic model.

Google tries hard to sell you on their auto-answers for emails ('smart reply'), wonder how those got trained...

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#264
post #153

Earlier quoted context omitted.

Do you really think the people asking ChatGPT to write their code can make that abstraction? The fact that the can't do this is the whole reason they have to use ChatGPT.

You can be an experienced developers with years building complex applications behind you and still find ChatGPT useful. I've found it useful for documenting individual methods or simply explaining my own/other's code or writing unit test methods or just using it to add boilerplate stuff that saves me an hour that I use elsewhere.

I think many people find ChatGPT useful specifically because they have years of experience building complex applications.

If you know exactly what you want to ask of it, and have the ability to evaluate and verify what it produces, it's incredible what you can get out of it. Sure it's nothing I couldn't have done otherwise... eventually. The productivity it enables is worth every cent.

Easily the best $20 I've spent in ages, they should have run with the initial idea of charging $42.

But holy moly anyone putting confidential information into it needs to stop

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#265
post #156
post #5

We block ChatGPT, as do most federal contractors. I think it’s a horrible exploit waiting to happen: - there’s no way they’re manually scrubbing out sensitive data so its bound to spill out from the training data when prompting the model - OpenAI is openly storing all this data they’re collecting to the extent that they’ve had several leaks now where people can see others’ conversations and data. We are one step away…

Does US intelligence have access to OpenAI data? Private organizations is one thing. But with all the dopes in government positions around the world, OpenAI logs would probably be a treasure trove for intelligence gathering.

USA has the patriot act and the cloud act to request the data from any USA company. Like AWS, Microsoft 365, Google…

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#266
post #5

We block ChatGPT, as do most federal contractors. I think it’s a horrible exploit waiting to happen: - there’s no way they’re manually scrubbing out sensitive data so its bound to spill out from the training data when prompting the model - OpenAI is openly storing all this data they’re collecting to the extent that they’ve had several leaks now where people can see others’ conversations and data. We are one step away…

So you block internet access for all employees? Cos anything you think is being pasted into ChatGPT is being pasted everywhere, whether its Google, Slack, Chrome Plugins, public Wifi.

Yes, these things are sometimes blocked in higher security workplaces… up to and including the public internet. Honestly airgapped systems are not all that uncommon anywhere that human life is at risk.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#267

Earlier quoted context omitted.

I can host llama's 7b model internally. It hallucinates more often than not and has a tendency to ramble, but dammit it's local and secure!

llama's 7b model internally is, for me, on a totally different level quality-wise. Even when explicitly instructed to not make up stuff and just say 'I don't know', it will still go ahead and ramble and invent things. When I tell it to only use the prompt data it will still invent, or just ignore the prompt data. It's not useful for production (i.e., to be exposed to 'regular' non AI users). ChatGPT, on the other han…

One problem with the current way these models are being trained is that they have no idea of what they're saying. It's just a recursive guess the next word type algorithm. I would not expect the confidence levels of any given fragment, let alone an average, to be a meaningful predictor of truth.

Also completely spitballing, I expect that a big chunk of OpenAI's 'secret sauce' is simple processing layers above and beyond the model. If you input gibberish to llama does it give you an output? If OpenAI is artificially tokenizing inputs (as opposed to just sending inputs straight to the software), it would both dramatically limit the input domain, thus improving output tuning, as well as give "it" the ability to say when it doesn't know something. I put "it" in quotes since that response would not becoming from the LLM, but from the preprocess tokenization system returning an error code in natural language.

I think there's some weak indirect evidence for this in the service itself, since incoherent inputs are instantly rejected, whereas even simple queries take dramatically longer to output even the first word. It's like the input is not even being sent to the LLM software for processing.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#268

Earlier quoted context omitted.

This really depends on the cost/benefit tradeoff for the entity in question. If using ChatGPT makes you X% more productive (shipping faster / lowers labor costs / etc), but comes with Y% risk of data leakage, is that worth it in expectation or not? I would argue that there definitely exist companies for which it's worth the tradeoff. By the way, OpenAI says they wont use data submitted through its API for model train…

To anyone who may be pasting code along the lines of 'convert this sql table schema into a [pydantic model|JSON Schema]' where you're pasting in the text, just ask it instead to write you a [python|go|bash|...] function that reads in a text file and 'converts an sql table schema to output x' or whatever. Related/not-related--great pandas docs replacement is another great+safe use-case. Point is, for a meaningful subs…

At first I was impressed by how easy it was to reach a data model with chatgpt, then I laughed as I tried to tweak it and use it. I realized it didn't really have any model concepts and was just using its various KB.

I am unsure if the so called AI can think in models but so far, not but still an impressive assisting tool if you take care of its limitations.

Another point where it lacks is in logic, my daughter has a lot of fun with the book "what is the name of this book?" but she was struggling with the "map of baal" explanation, her the answer was a certain map, yet the book had another answer, I had a third one as I interpreted a proposition. I never got an answer without a contradiction in chatgpt reasoning, and the book had been mistranslated to French so one of its propositions was changed (C, both A and B were knaves) but not the answer.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#269

Earlier quoted context omitted.

Did you read the next paragraph? > When you use our non-API consumer services ChatGPT or DALL-E, we may use the data you provide us to improve our models.

I definitely did not correctly read that. Thanks for the clarification. Totally misread the 'our API' bit! It's also in the FAQ: https://help.openai.com/en/articles/6783457-chatgpt-general-... > Will you use my conversations for training? > Yes. Your conversations may be reviewed by our AI trainers to improve our systems.

As the saying goes, if the product is free, then you are the product.

Re: Employees are feeding sensitive data to ChatGPT, raising security fears

#270
post #153

Earlier quoted context omitted.

To anyone who may be pasting code along the lines of 'convert this sql table schema into a [pydantic model|JSON Schema]' where you're pasting in the text, just ask it instead to write you a [python|go|bash|...] function that reads in a text file and 'converts an sql table schema to output x' or whatever. Related/not-related--great pandas docs replacement is another great+safe use-case. Point is, for a meaningful subs…

Do you really think the people asking ChatGPT to write their code can make that abstraction? The fact that the can't do this is the whole reason they have to use ChatGPT.

Precisely because I can abstract it is why I use ChatGPT. It can do the boring, tedious, repetitive stuff instead of me and has shown me the joy of using programming to solve ACTUAL problems yet again, instead of having to spend hours on unimportant problems like "how do I do X with library Y".
Post reply on HN