Earlier quoted context omitted.
I thought home use is whatever fits in 24GB (a single 3090 GPU, which is pretty affordable), not 8 or 16. 30B models fit.
While some home users do indeed have 24GB of vram, the fact is a 4090 costs $1700 Such models will never top the number of downloads charts, or the community hype, as there’s just loads more people who can use the smaller models. And if you can afford one 4090 you can probably afford two.
Llama 3.1
271–279 of 279 posts
Re: Llama 3.1
#272Can this Llama process ~1GB of custom XML data? And answer queries like: Give all which refer to which refer to an Indo-European .
Re: Llama 3.1
#273This 405B seriously need quantization solution like 1.625 bpw ternary packing for BitNet b1.58 https://github.com/ggerganov/llama.cpp/pull/8151
The perplexity per parameter is higher and the delta grows as it scales.
Not per bit, but per parameter.
Why this is happening really needs more attention and more consideration for pretrained model development right now.
A sleeping giant of a difference in a space where even marginal gains make headlines.
Re: Llama 3.1
#274Re: Llama 3.1
#275You can chat with these new models at ultra-low latency at groq.com. 8B and 70B API access is available at console.groq.com. 405B API access for select customers only – GA and 3rd party speed benchmarks soon. If you want to learn more, there is a writeup at https://wow.groq.com/now-available-on-groq-the-largest-and-m... . (disclaimer, I am a Groq employee)
We also added Llama 3.1 405B to our VSCode copilot extension for anyone to try coding with it. Free trial gets you 50 messages, no credit card required - https://double.bot (disclaimer, I am the co-founder)
Re: Llama 3.1
#276Earlier quoted context omitted.
Super cool, though sadly 405b will be outside most personal usage without cloud providers which sorta defeats the purpose of opensource to some extent atleast sadly, because .. nvidia's rampup of consumer VRAM is glacial
Great for Groq whos already hosting it but at what cost I guess.
Re: Llama 3.1
#277Earlier quoted context omitted.
> at home with the right hardware Where the right hardware is 10x4090s even at 4 bits quantization. I'm hoping we'll see these models get smaller, but the GPT-4-competitive one isn't really accessible for home use yet. Still amazing that it's available at all, of course!
It's hardly cheap starting at about $10k of hardware, but another potential option appears to be using Exo to spread the model across a few MBPs or Mac Studios: https://x.com/exolabs_/status/1814913116704288870
Re: Llama 3.1
#278You can chat with these new models at ultra-low latency at groq.com. 8B and 70B API access is available at console.groq.com. 405B API access for select customers only – GA and 3rd party speed benchmarks soon. If you want to learn more, there is a writeup at https://wow.groq.com/now-available-on-groq-the-largest-and-m... . (disclaimer, I am a Groq employee)