Llama 3.1
141–150 of 279 posts
Re: Llama 3.1
#142Re: Llama 3.1
#143I'm glad to see the nice incremental gains on the benchmarks for the 8B and 70B models as well.
Some of those benchmarks show quite significant gains. Going from Llama-3 to Llama-3.1, MMLU scores for 8B are up from 65.3 to 73.0, and 70B are up from 80.9 to 86.0. These scores should always be taken with a grain of salt, but this is encouraging. 405B is hopelessly out of reach for running in a homelab without spending thousands of dollars. For most people wanting to try out the 405B model, the best option is to r…
Re: Llama 3.1
#144Earlier quoted context omitted.
how is this even useful? no one can run it.
If you want to run the 405B model without spending thousands of dollars on dedicated hardware, you rent compute from a datacenter. Meta lists AWS, Google and Microsoft among others as cloud partners. But also check out the 8B and 70B Llama-3.1 models which show improved benchmarks over the Llama-3 models released in April.
Re: Llama 3.1
#145Earlier quoted context omitted.
Unsure if anyone has specific hardware benchmarks for the 405b model yet, since it's so new, but elsewhere in this thread I outlined a build that'd probably be capable of running a quantized version of Llama 3.1 405b for roughly $10k. The $10k figure is likely roughly the minimum amount of money/hardware that you'd need to run the model at acceptable speeds, as anything less requires you to compromise heavily on GPU…
Energy costs are an important factor here too. While Quadro cards are much more expensive upfront (higher $/VRAM), they are cheaper over time (lower Watts/Token). Offsetting the energy expense of a 3090/4090/5090 build via solar complicates this calculation but generally speaking can be a "reasonable" way of justifying this much hardware running in a homelab. I would be curious to see relative failure rates over time…
I don't have any actual figures to back this up, but my gut tells me that the fact that enterprise GPUs are an order of magnitude (at least) more expensive than, say a, 3090, means that the payback period of them has got to be pretty long. I also wonder whether setting the max power on a 3090 to a lower than default value (as I suggest in my other post) has a significant effect on the average W/token.
Re: Llama 3.1
#146Earlier quoted context omitted.
Super cool, though sadly 405b will be outside most personal usage without cloud providers which sorta defeats the purpose of opensource to some extent atleast sadly, because .. nvidia's rampup of consumer VRAM is glacial
You don't need a model of this scale for personal use. Llama 3.1 8B can easily run on your laptop right now. The 70B model can run on a pair of 4090s.
This is good enough for a lot of use cases... on a laptop. An expensive laptop, but hardware only gets better and cheaper over time.
Re: Llama 3.1
#147Open Source AI Is the Path Forward - Mark Zuckerberg https://about.fb.com/news/2024/07/open-source-ai-is-the-path...
So are they actually making the models open now or are they staying the course with "kind of open" as they have done for LLaMA 1, 2, and 3 [1]? [1]: https://opensource.org/blog/metas-llama-2-license-is-not-ope... As I have stated time and again, it is perfectly fine for them to slap on whatever license they see fit as it is their work. But it would be nice if they used appropriate terms so as not to disrupt the disco…
It's "a Google and Apple can't use this model in production" clause that frankly we can all be relatively okay with.
Re: Llama 3.1
#148Re: Llama 3.1
#149Earlier quoted context omitted.
So are they actually making the models open now or are they staying the course with "kind of open" as they have done for LLaMA 1, 2, and 3 [1]? [1]: https://opensource.org/blog/metas-llama-2-license-is-not-ope... As I have stated time and again, it is perfectly fine for them to slap on whatever license they see fit as it is their work. But it would be nice if they used appropriate terms so as not to disrupt the disco…
> specifically, it puts restrictions on commercial use for some users (paragraph 2) and also restricts the use of the model and software for certain purposes (the Acceptable Use Policy) It's "a Google and Apple can't use this model in production" clause that frankly we can all be relatively okay with.
Re: Llama 3.1
#150Nice, someone donate me a few 4090s :(
maybe someone will figure out some ways to prune/ quantize it a huge amount ;-; edit: If the AI bubble pops we will be swimming in GPUs... but no new models.