Live data from Hacker News

Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support

github.com

171–180 of 298 posts

Re: Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support

#171

Earlier quoted context omitted.

LLaMA-65B fits in 32GB of VRAM using state of the art GPTQ quantization with no output performance loss. https://github.com/qwopqwop200/GPTQ-for-LLaMa

So if I'm reading this right, 65B at 4bit would consume around 20GB of VRAM and ~130GB of system RAM?

LLaMA it doesn't require any system RAM to run.

It requires some very minimal system RAM to load the model into VRAM and to compile the 4bit quantized weights.

But if you use pre-quantized weights (get them from HuggingFace or a friend) then all you really need is ~32GB of VRAM and maybe around 2GB of system RAM for 65B. (It's 30B which needs 20GB of VRAM.)

Re: Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support

#172
post #86

Earlier quoted context omitted.

PCIe devices like GPUs can access system memory. Integrated GPUs also access system memory via the same bus as the CPU. It’s not really a new technique. Apple just shipped a highly integrated unit with large memory bandwidth.

That's a pretty important just. They chose to go down this path at this time and shipped, now I can run the 64B Llama on a widely available $3,999 prosumer device. How much would a PC that can do that currently cost me and can I have it by tomorrow?

>How much would a PC that can do that currently cost me and can I have it by tomorrow?

At the moment, seems like Apple has an edge here. On PC for single GPU you need an NVIDIA A40, which used prices for is about $2500, and not at retail stores.

If you don't mind having two GPUs then two $800 3090 GPUs works, but that's a workstation build you'll have to order from Puget or something. That's probably faster than Apple.

My gut instinct is that there's some low hanging fruit here and in the next couple weeks 64B Llama will run comparably or faster on any PC with a single 4090/3090 and 64 or 128 GB of system memory. But probably not any PC laptops that aren't 17 inch beasts, Apple will keep that advantage.

Re: Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support

#174
post #86

Earlier quoted context omitted.

That's a pretty important just. They chose to go down this path at this time and shipped, now I can run the 64B Llama on a widely available $3,999 prosumer device. How much would a PC that can do that currently cost me and can I have it by tomorrow?

You can buy a prebuilt pc with a 4090 for less - which is significantly more powerful but still in the 3xxx$. You could go cheaper with a 3090 which has the same vram and it's just slower. I think the best combo is a serious Nvidia pc for AI + a cheap MacBook air for portability.

You need two 3090s or 4090s to fit 65B even in 4bit. It's a big one.

That said, if you're fine with slower speeds then two P40s could get the job done for $150 each. (Not sure how much slower this would go though.)

Re: Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support

#175
post #60

> Currently, only LLaMA-7B is supported since I haven't figured out how to merge the tensors of the bigger models. However, in theory, you should be able to run 65B on a 64GB MacBook Suddenly the choices Apple made with its silicon are looking like pure genius as there will be a lot of apps using this that are essentially exclusive to their platform. Even with the egregious ram pricing. With a lot of fine tuning if y…

Wonder why AMD/Intel/Nvidia haven’t invented some sort of device that allows the processor and graphics to share memory like Apple has done.

They have, see eg this from 10 years ago: https://www.tomshardware.com/news/AMD-HSA-hUMA-APU,22324.htm...

Implementing the software support and getting operating systems to play along and fragmentation between GPU vendors, as always with GPUs on x86, have been the problems. From all accounts it's been working reasonably well on the consoles though.

Also chicken-and-egg: low GPU compute usage uptake outside of games has meant it's not improved lately.

Re: Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support

#176

Earlier quoted context omitted.

You're likely right that applying RLHF (+ fine-tuning with instructions) to LLaMA 7b would produce results similar to ChatGPT, but I think you're implying that that would be feasible today. RLHF requires a large amount of human feedback data and IIRC there's no open data set for that right now.

There are open datasets (see the chatllama harness project and its references). You can of course also cross train it using actual ChatGPT.

Is there something I'm missing? ChatLlama doesn't reference any human feedback datasets.

> You can of course also cross train it using actual ChatGPT.

You mean train it on ChatGPT's output? That's against OpenAI's terms of service.

Re: Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support

#177
post #104

A quick survey of the thread seems to indicate the 7b parameter LLaMA model does about 20 tokens per second (~4 words per second) on a base model M1 Pro, by taking advantage of Apple Silicon’s Neural Engine. Note that the latest model iPhones ship with a Neural Engine of similar performance to latest model M-series MacBooks (both iPhone 14 Pro and M1 MacBook Pro claim 15.8 teraflops on their respective neural engines…

>20 tokens per second (~4 words per second)

How can there be 5 tokens per word, when they have more than half the vocabulary as GPT-2/3 which has 1.3 tokens per word?

I would have guessed more like 1.5 tokens per word.

Re: Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support

#178
post #104

A quick survey of the thread seems to indicate the 7b parameter LLaMA model does about 20 tokens per second (~4 words per second) on a base model M1 Pro, by taking advantage of Apple Silicon’s Neural Engine. Note that the latest model iPhones ship with a Neural Engine of similar performance to latest model M-series MacBooks (both iPhone 14 Pro and M1 MacBook Pro claim 15.8 teraflops on their respective neural engines…

>20 tokens per second (~4 words per second) How can there be 5 tokens per word, when they have more than half the vocabulary as GPT-2/3 which has 1.3 tokens per word? I would have guessed more like 1.5 tokens per word.

Oh, it’s probably higher than four words per second, then. I assumed tokens was characters and used the standard “there are five characters in a word” rule of thumb.

Re: Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support

#179
I can confirm that this (7B) runs nicely on a 24GB MacBook Air M2. The output of my initial test was definitely a bit different than ggreganov's example!

The first man on the moon was 39 years old on July 16, 1969. July 16th is the 198th day of the year (199th in leap years) in the Gregorian calendar. There are 168 days remaining until the end of the year. 1561 – France is divided into 2535 circles (French: cercles) for fiscal purposes. 1582 – Pope Gregory XIII, through a papal bull, establishes the Gregorian calendar (Old Style and New Style dates). 1

Re: Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support

#180

Earlier quoted context omitted.

There are open datasets (see the chatllama harness project and its references). You can of course also cross train it using actual ChatGPT.

Is there something I'm missing? ChatLlama doesn't reference any human feedback datasets. > You can of course also cross train it using actual ChatGPT. You mean train it on ChatGPT's output? That's against OpenAI's terms of service.

You’re missing something. Both SHP (https://huggingface.co/datasets/stanfordnlp/SHP) and OpenAssistant datasets are referenced.

And the TOS violation might be the case, the project nevertheless has a mode to use OpenAI in the fine tuning steps.

Post reply on HN