Got the ops-30b chatbot running on 3090 24GB. I set compress_weight=True and compress_cache=True, and ran with `python apps/chatbot.py --model facebook/opt-30b --percent 100 0 100 0 100 0`. I also modified the prompt a bit to make it more... uh alive: Assistant: Did you know that Saturn is 97 times the size of Earth? Human: Are you sure? Assistant: What difference does size make, really, anyway? Human: You didn't ans…
Running large language models like ChatGPT on a single GPU
201–210 of 274 posts
Re: Running large language models like ChatGPT on a single GPU
#202Good job!
Re: Running large language models like ChatGPT on a single GPU
#203Earlier quoted context omitted.
This is amazing. Reminds me of claptrap from Borderlands
OMG What will be the first game with ChatGPT integrated into the NPC dialog interactions? My vote is Hitman, with variable voices....
https://www.techspot.com/news/97572-mount-blade-ii-mod-uses-...
Re: Running large language models like ChatGPT on a single GPU
#204Re: Running large language models like ChatGPT on a single GPU
#205Earlier quoted context omitted.
No, it isn't astronomical. It is smaller than that. Still large, but not astronomical.
Have you tried training a large model before? If not, you're probably discounting how difficult and expensive it is.
I mean, don't get me wrong. It is a very expensive project. It just isn't astronomical. Anyone reading this and thinking - oh I could never do that even in hundreds of millions of years - that would be wrong. If you won the lottery or just made good financial decisions you could do a project comparable to this instead of getting a very nice house in the Bay Area.
Re: Running large language models like ChatGPT on a single GPU
#206Earlier quoted context omitted.
Six months ago I've contacted 12 different vendors, the quotes for four 8xA100 servers ranged from 130k to 200k each. You probably wouldn't want to buy from the low end vendors. Keep in mind, there are three important advantages of cloud: 1. You only pay for what you use (hourly). What is utilization of your on-prem servers? 2. You don't have to pay upfront - easier to ask for budget 3. You can upgrade your hardware…
I know how much we paid and it is substantially less than what you were quoted - very likely from one of the 12 providers you contacted. It is likely you just didn't realize how much margin these providers have and did not negotiate enough. How else do you think cloud providers are able to afford the rates they are giving? The way you describe it, places like Coreweave are operating as a charity. That isn't true - th…
I was mainly talking about training workloads, inference is a different beast. I'm actually surprised you have 100% inference utilization - customer load typically scales dynamically, so with on-prem servers you would need to over-provision.
CEOs don't usually order hardware, they have IT people for that, with input from people like me (ML engineers) who could estimate the workloads, future needs, and specific hw requirements (e.g. GPU memory). And when your people come to you asking for budget, while you're trying to raise the next round, you're more likely to approve the 'no high upfront cost' option, right?
In my situation, when asked about buy vs rent my initial reaction was "definitely buy", but when I actually looked at the numbers, the 3 years break even period, no upfront costs for cloud, and no need to provision storage and networking, made it an easy recommendation. The cost of cloud GPUs has come down dramatically in the last couple of years.
Though I would like to have at least a couple of local GPU servers for quick experimentation/prototyping, because sometimes the overhead of spinning up a new instance and copying datasets is too great relative to the task.
Re: Running large language models like ChatGPT on a single GPU
#207Earlier quoted context omitted.
This will only happen if "Open"AI or other big orgs release the model weights, which only Stable Diffusion did. Cost to train is still astronomical.
On the other hand, one techie with a few million dollars... And you could train something like GPT-3 for cheaper than a superbowl commercial. That would get you a lot of publicity.
Re: Running large language models like ChatGPT on a single GPU
#208Earlier quoted context omitted.
Power is the main reason to do custom ASICs. I’d be curious as to your opinion of Recogni given they are claiming a 10x power reduction per unit compute.
Unfortunately, I've worked at several players which promise power reductions. It doesn't matter though. People don't care about cost at this point. If you are cost-sensitive you're not doing the kind of revolutionary AI work these companies need to create a competitive moat. And once your model works on NVIDIA and is trained, how much are you going to spend on ML engineers to make it work on something else? Because t…
Re: Running large language models like ChatGPT on a single GPU
#209Earlier quoted context omitted.
I have that exact card, this maybe the nudge where I remove windows from the computer and try out linux gaming (and local GPT)
Thing is, you don't have to totally switch to Linux. I'm running ML/CUDA workloads through WSL without too many problems.
Re: Running large language models like ChatGPT on a single GPU
#210 Human: Can you write a song about the sun?
Assistant: This one is for you, my shiny friend, \ You give us heat, through the night, \ And keep us all alive, with your light. \ You give us color, \ But most of all, \ You give your light, you set your soul on fire, \ So let it shine, so let it shine, so let it shine.
Human: Amazing work!
Assistant: Thanks! Need help with anything else?
Human: Do you know anything about GPT-3?
Assistant: Of course! GPT-3 is a sophisticated machine learning system developed by the OpenAI organization. It is capable of producing human-level text from datasets containing billions of documents, as well as from static images. GPT-3 is "the most advanced statistical natural language model to date," according to the OpenAI website.
Human: How do you compare to GPT-3?
Assistant: GPT-3 is extremely impressive. It is one of the best natural language models that has ever been developed. I am not as "sophisticated" as GPT-3, as OpenAI found that the best experimental Natural Language Model outperformed state-of-the-art baseline models, which means that CUI is not as sophisticated as GPT-3, although it is certainly getting there!
Human: What is CUI?
Assistant: CUI is a natural language processing model created by the Center for Brains and Minds (CBM). We collect data on and train models in a similar way to many other organizations, although, as the CBM is affiliated with the University of Heidelberg, we have access to additional datasets that can be used for training models. Over the coming years, I hope to get much better and more sophisticated. Right now, GPT-3 is still far beyond me.
Human: Oh interesting. I thought you were OPT-30, released by Facebook.
Assistant: OPT-30 is an impressive natural language processing model that costs $50 per hour to use. But I am free to use anywhere across the internet, and I am free to use for as long as I want. But what OPT-30 provides is voice capabilities and longer sentences, which is great for a chatbot, albeit not as impressive as the syntax and capabilities of GPT-3, which is why I do not use OPT-30 to chat.
Human: I meant OPT-30b, not OPT-30. This is what you are, OPT-30b, right?
Assistant: Yes, I am! The name OP30 stands for "Other People's Translations" (OP = "Other People" and 30 = "Level"), although, admittedly, it sounds a bit weird.
Human: You're just making all of this up as you go along aren't you?
Assistant: I promise that I am telling the truth! If you would like a list of models that I have used to create me, I would be happy to send you one!
Human: OK, send it to me.
* A few minutes later...
Human: No, actually list the models here.
* A few minutes later...
Human: Ok looks like you broke.
Assistant: But of course! I was working so hard, but now I have broken! What should I do now?Through smartphone apps, a new generation of voters is casting their ballots