Live data from Hacker News

Meta Llama 3

llama.meta.com

551–560 of 965 posts

Re: Meta Llama 3

#551

Earlier quoted context omitted.

It's a very plausible rumor, but it is misleading in this context, because the rumor also states that it's a mixture of experts model with 8 experts, suggesting that most (perhaps as many as 7/8) of those weights are unused by any particular inference pass. That might suggest that GPT-4 should be thought of as something like a 250B model. But there's also some selection for the remaining 1/8 of weights that are used…

What is the reason for settling on 7/8 experts for mixture of experts? Has there been any serious evaluation of what would be a good MoE split?

A 19" server chassis is wide enough for 8 vertically mounted GPUs next to each other, with just enough space left for the power supplies. Consequently 8 GPUs is a common and cost efficient configuration in servers.

Everyone seems to put each expert on a different GPU in training and inference, so that's how you get to 8 experts, or 7 if you want to put the router on its own GPU too.

You could also do multiples of 8. But from my limited understanding it seems like more experts don't perform better. The main advantage of MoE is the ability to split the model into parts that don't talk to each other, and run these parts in different GPUs or different machines.

Re: Meta Llama 3

#552

Earlier quoted context omitted.

For the still training 400B: Llama 3 GPT 4(Published) BBH 85.3 83.1 MMLU 86.1 86.4 DROP 83.5 80.9 GSM8K 94.1 92.0 MATH 57.8 52.9 HumEv 84.1 74.4 Although it should be noted that the API numbers were generally better than published numbers for GPT4. [1]: https://deepmind.google/technologies/gemini/

Those numbers are for the original GPT-4 (Mar 2023). Current GPT-4-Turbo (Apr 2024) is better: Llama 3 GPT-4 GPT-4-Turbo* (Apr 2024) MMLU 86.1 86.4 86.7 DROP 83.5 80.9 86.0 MATH 57.8 52.9 73.4 HumEv 84.1 74.4 88.2 *using API prompt: https://github.com/openai/simple-evals

I find it somewhat interesting that there is a common perception about GPT-4 at release being actually smart, but that it got gradually nerfed for speed with turbo, which is better tuned but doesn't exhibit intelligence like the original.

There were times when I felt that too, but nowadays I predominantly use turbo. It's probably because turbo is faster and cheaper, but in lmsys turbo has 100 elo higher than original, so by and large people simply find turbo to be....better?

Nevertheless, I do wonder if not just in benchmarks but in how people use LLMs, intelligence is somewhat under utilised, or possibly offset by other qualities.

Re: Meta Llama 3

#553
post #189
post #184

If anyone is looking to try 7B locally really quick, we have just added it to Msty. [1]: https://msty.app

From the faq > Does Msty support GPUs? > Yes on MacOS. On Windows* only Nvidia GPU cards are supported; AMD GPUs will be supported soon. Do you support GPUs on linux? Your downloads with windows are also annotated with CPU/CPU + GPU, but your linux ones aren't. Does that imply they are CPU only?

Yes, if CUDA drivers are installed it should pick it up.

Re: Meta Llama 3

#554
"You’ll also soon be able to test multimodal Meta AI on our Ray-Ban Meta smart glasses."

Now this is interesting. I've been thinking for some time now that traditional computer/smartphone interfaces are on the way out for all but a few niche applications.

Instead, everyone will have their own AI assistant, which you'll interact with naturally the same way as you interact with other people. Need something visual? Just ask for the latest stock graph for MSFT for example.

We'll still need traditional interfaces for some things like programming, industrial control systems etc...

Re: Meta Llama 3

#555

Earlier quoted context omitted.

Having an engineering mindset is not the same as never making mistakes (or never being too early to the market). The only way you won’t make those mistakes and keep a perfect record is if you never do anything major or step out of the comfort zone. If Apple didn’t try and fail with Newton[0] (which was too early to the market for many reasons, both tech-related and not), we might’ve not had iPhone today. The engineer…

His engineering mindset made him blind to the fact the metaverse was a product that nobody wanted or needed. In one of the Fridman interviews, he goes on and on about all the cool technical challenges involved in making the metaverse work. But when Fridman asked him what he likes to do in his spare time, it was all things that you could precisely not do in the metaverse. It was baffling to me that he failed to connec…

Let’s be honest VR is about the porn. I’d it’s successful at that Zuck will make his billions.

Re: Meta Llama 3

#558
post #552

Earlier quoted context omitted.

Those numbers are for the original GPT-4 (Mar 2023). Current GPT-4-Turbo (Apr 2024) is better: Llama 3 GPT-4 GPT-4-Turbo* (Apr 2024) MMLU 86.1 86.4 86.7 DROP 83.5 80.9 86.0 MATH 57.8 52.9 73.4 HumEv 84.1 74.4 88.2 *using API prompt: https://github.com/openai/simple-evals

I find it somewhat interesting that there is a common perception about GPT-4 at release being actually smart, but that it got gradually nerfed for speed with turbo, which is better tuned but doesn't exhibit intelligence like the original. There were times when I felt that too, but nowadays I predominantly use turbo. It's probably because turbo is faster and cheaper, but in lmsys turbo has 100 elo higher than original…

Have you tried Claude 3 Opus? I've been using that predominantly since release and find it's "smarts" as or better than my experience with GPT-4 (pre turbo).

Re: Meta Llama 3

#559
post #469

Earlier quoted context omitted.

> Meta AI isn't available yet in your country Where is it available? I got this in Norway.

This is so frustrating. Why don't they just make it available everywhere?

I'm always glad at these rare moments when EU or American people can get a glimpse of a life outside the first world countries.

Re: Meta Llama 3

#560

Earlier quoted context omitted.

Would you then say that in general Open Source doesn't matter for almost everyone? Most people running Linux aren't serving 700 million customers or operating military killbots with it after all.

> in general Open Source doesn't matter for almost everyone? Most of the qualities that come with open source (which also come with llama 3), matter a lot. But no, it is not a binary, yes or no thing, where something is either open source and useful or not. Instead, there is a very wide spectrum is licensing agreements. And even if something does not fit the very specific and exact definition of open source, it can s…

If I build a train, put it into service, and say to the passengers “this has 99.9% of the required parts from the design”, would you ride on that train? Would you consider that train 99.9% as good at being a train? Or is it all-or-nothing?

I don’t necessarily disagree with your point about there still being value in mostly-open software, but I want to challenge your notion that you still get most of the benefit. I think it being less than 100% open does significantly decay the value, since now you will always feel uneasy adopting these models, especially into an older existing company.

You can imagine a big legacy bank having no problem adopting MIT code in their tech. But something with an esoteric license? Even if it’s probably fine to use? It’s a giant barrier to their adoption, due to the risk to their business.

That’s also not to say I’m taking it for granted. I’m incredibly thankful that this exists, and that I can download it and use it personally without worry. And the huge advancement that we’re getting, and the public is able to benefit from. But it’s still not the same as true 100% open licensing.

Post reply on HN