Any more details? I'm guessing they're leveraging the NPU in the pixel?
LLaMa running at 5 tokens/second on a Pixel 6
11–20 of 79 posts
Re: LLaMa running at 5 tokens/second on a Pixel 6
#12Any more details? I'm guessing they're leveraging the NPU in the pixel?
Re: LLaMa running at 5 tokens/second on a Pixel 6
#13I'm waiting until it runs on my C64...
Re: LLaMa running at 5 tokens/second on a Pixel 6
#14Did anyone get this to run on an iPhone or in a browser yet?
Re: LLaMa running at 5 tokens/second on a Pixel 6
#15Re: LLaMa running at 5 tokens/second on a Pixel 6
#16Does this in theory mean it should be relatively easy to port to coral TPU?
So maybe if you implement the ggml.c with tensorflow/libcoral - you'd have a chance.
Re: LLaMa running at 5 tokens/second on a Pixel 6
#17Earlier quoted context omitted.
From the video output seems fine. But if it is a trimmed version, it is wong to call it LLaMa.
It's nonsensical, celeb announces they're going to rehab and notes it (?) is an issue affecting all women, at least, earlier today (??), they also noted it wasn't drugs or alcohol this time, but, a life (???)
Re: LLaMa running at 5 tokens/second on a Pixel 6
#18Re: LLaMa running at 5 tokens/second on a Pixel 6
#19It is not really llama, it is llama quantized to 4bit. Not even the quality of original 7B. I could also quantize it to 1 bit and claim it runs on my RPI3.
Re: LLaMa running at 5 tokens/second on a Pixel 6
#20It is not really llama, it is llama quantized to 4bit. Not even the quality of original 7B. I could also quantize it to 1 bit and claim it runs on my RPI3.