Earlier quoted context omitted.
If this is a "wake up call" - then your legal team needs immediate education. First - there is this - https://openai.com/policies/how-your-data-is-used-to-improve... (linked from the Navier Stokes writeup) I don't know how much more clearly they can write: > When you use our services for individuals such as ChatGPT, Sora, or Operator, we may use your content to train our models. One of the key selling tactics that co…
The training on user data only applies to free accounts - paid and Enterprise accounts guarantee data is not used for training. Plenty of Enterprises use the APIs directly - that's just plain misinformation
There's a reason why people spend more $$$ with Data Bricks, Palantir, AWS Bedrock etc.. and don't even consider using Anthropic or OpenAI APIs directly - it's because those guarantees provide very little in the way of data-discovery, audit requirements, or liquidated damages should it ever be discovered there was data leakage.
At least with these other companies, while the LD is likewise not great (typically limited to the amount of money you paid them) - you at least have some data-governance guarantees around running on dedicated hardware - no multi-tenancy, no third-party access outside of the AWS operators who keep the HW running - but are very much not in the business of looking at your data.
I think this is mostly a function of what's at risk - when company valuations get into the 10s of billions of dollars, the risk of IP leaking into what could be seen as competitive companies (OpenAI/Anthropic would be happy to take over the world - I don't sense that AWS or Azure, are as ruthless in stepping on their customers business, unless of course they are a SAAS provider) is just too significant a liability to take - particularly when you can de-risk.