Earlier quoted context omitted.
OpenAI says they offer a BAA but I haven’t heard of anyone actually being able to get to someone who could put one together. You can get a BAA through Azure’s OpenAI service though, I believe the details are located in this document: https://azure.microsoft.com/en-us/resources/microsoft-azure-...
Pretty sure I know an org that has baa with OpenAI directly. Agree that azure is more straightforward (and what I did at my startup).
ChatGPT Enterprise
481–490 of 532 posts
Re: ChatGPT Enterprise
#482Earlier quoted context omitted.
In your opinion, is it an either or scenario? Or would fine-tuning on docs + RAG be even more powerful?
I've been wondering this myself lately. After using RAG with pgvector for the last few months with temperature 0, it's been pretty great with very little hallucination. The small context window is the limiting factor. In principle, I don't see the difference between a bunch of fine-tuned prompts along the lines of "here is another context section: ", which is the same as what it looks like in a RAG prompt anyway. May…
Other considerations: (A) would you fine-tune daily? weekly? as data changes? (B) Cost and availability of GPUs (there's a current shortage)
My experience is that RAG is the way to go, at least right now.
But you have to make sure your retrieval engine work optimally: getting the very most relevant pieces of text from your data: (1) using a good chunking strategy that's better than arbitrary 1K or 2K chars (2) using a good embedding model (3) Using hybrid search, and a few other things like that.
Certainly the availability of longer sequence models is a big help
Sharing this relevant discussion from LinkedIn: https://www.linkedin.com/feed/update/urn:li:activity:7101638...
Re: ChatGPT Enterprise
#483Earlier quoted context omitted.
> If your company illegally treads on my IP, I don't care if your employees used LLM's or not; I will sue. In the US the saying is "This case will come down to who has the last dime and it won't be you". > By the way, I don't have to win a lawsuit to get some justice; I'll make discovery hurt. Even for discovery a common tactic deployed for discovery is to "bury" the other party in discovery materials to the point wh…
Perhaps, but I have a plan to take that burying and make it hurt. And yes, I know that money wins in legal battles. That's why I am focusing on discovery before I don't have any.
When they send you several thousand pages (minimum) of random documents you're looking for a smoking gun that's a needle in a haystack.
That takes money and time that very few parties have, coming back to the last dime expression.
I don't know if you've ever deployed your strategy successfully but needless to say I wouldn't count on it as a viable path to protect your IP against companies valued in the tens of billions of dollars.
Re: ChatGPT Enterprise
#484Explicitly calling out that they are not going to train on enterprise's data and SOC2 compliance is going to put a lot of the enterprises at ease and embrace ChatGPT in their business processes. From our discussions with enterprises (trying to sell our LLM apps platform), we quickly learned how sensitive enterprises are when it comes to sharing their data. In many of these organizations, employees are already pasting…
The SOC2 framework is complex and compliance can be expensive. This can lead organizations to focus on ticking the boxes rather than implementing meaningful security controls.
SOC2 is not a good universal metric for understanding an organization's security culture. It's frightening that this is the best we have for now.
Re: ChatGPT Enterprise
#485Earlier quoted context omitted.
> The ChatGPT model has violated pretty much all open source licenses Are you claiming this because they used copyrighted material as training data? If so, I think you're starting from the wrong point. Please correct me if I'm wrong, but last I heard using copyrighted data is pretty murky waters legally and they're operating in a gray area. Additionally, I don't think many open source licenses explicitly forbid using…
> Are you claiming this because they used copyrighted material as training data? If so, I think you're starting from the wrong point. All open source license comes under copyright law. It means if they violate the OSS license, the license is void and the tech/material becomes copyright protected. So yes, it would mean that it is trained on copyrighted material. > Additionally, I don't think many open source licenses…
Re: ChatGPT Enterprise
#486Earlier quoted context omitted.
Why do people bring this up? People are not LLMs and the issues are not the same.
Why are the issues not the same? Are you privileging meat over silicon?
They are not the same because an LLM is a construct. It is not a living entity with agency, motive, and all the things the law was intended for.
We will see new law as this tech develops.
For an analogy, many people call infringement theft and they are wrong to do so.
They will focus on the someone getting something without having followed the right process part while ignoring the equally important someone else being denied the use of, or loss of property part.
The former is an element in common between theft and infringement. And it is compelling!
But, the real meat in theft is all about people losing property! And that is not common at all.
This AI thing is similar. The common elements are super compelling.
But it just won't be about that in the end. It will be all about the details unique to AI code.
Re: ChatGPT Enterprise
#487Earlier quoted context omitted.
I've been experiencing carpal tunnel on and off for a couple of weeks now. I can tell you that reading through some code generated by "insert llm x" is substantially less painful than writing all of it by my own hand. Especially if you start understanding how to refine your prompts to the point where you use a single thread for project management and use that thread to generate prompts for other threads. Not all valu…
Take sick leave
A guy could wonder why so many of us do not use those answers.
Could it be the details complicate things just enough to take the easy answer off the table?
Perhaps it is just me. What say you?
Re: ChatGPT Enterprise
#488Earlier quoted context omitted.
Are you really asserting that these models aren't learning? What definition of learning are you using?
Don't know if they are, and don't really care either and I don't care to anthropomorphize circuitry to the extent that AI proponents tend to, especially. Humans and Computers are 2 wholly separate entities, and there's 0 reason for us to conflate the two. I don't care if another human looks at my code and straight up copies/pastes it, I care very much if an entity backed by a megacorp like Micro$oft does the same, en…
However, on the other hand we also have the scale at which they learn, which kind of makes every individual source line of code they learn from pretty unimportant. Learning at this scale is statistical process, and in most cases individual source snippets diminish in the aggregation of millions of others.
Or to put it the other way round, the actual value lies in the effort of collecting the samples, training the models, creating the software required for the whole process, putting everything into a good product and selling it. Again, in my mind, the importance of every individual source repo is too small at this scale to care about their license.
Re: ChatGPT Enterprise
#489Earlier quoted context omitted.
There are great tools that do this already in a support-multiple-ecosystems kind of way! I'm actually the CEO of one of those tools: Credal.ai - which lets you point-and-click connect accounts like O365, Google Workspace, Slack, Confluence, e.t.c, and then you can use OpenAI, Anthropic etc to chat/slack/teams/build apps drawing on that contextual knowledge: all in a SOC 2 compliant way. It does use a Retrieval-Augmen…
What are the limitations on adding documents to your system? Your website doesn't particularly highlight that feature set, which it probably should if you support it!
Re: ChatGPT Enterprise
#490Clicked on ChatGPT / Compare ChatGPT plans / Enterprise ... > Contact sales Oops. Scary. I'm missing the Teams plan: transparent pricing with a common admin console for our team. Yes, fast GPT-4, 32k context, templates, API credits... they're all very nice-to-haves, but just the common company console would be crucial for onboarding and scaling-up our team and needs without the big-bang "enter-pricey" stuff.
Any "Contact sales" stuff has just been an instant "no" at any company I've ever worked at, because that always means that the numbers are always too high to include in the budget unless it's a directive coming down directly from the top.