Live data from Hacker News

Meta Llama 3

llama.meta.com

171–180 of 965 posts

Re: Meta Llama 3

#171
post #168
post #161

Earlier quoted context omitted.

[flagged]

What a silly, provocative comparison. China is a suppressive state that strives to control its citizens while the EU privacy protection laws are put in place to protect citizens. If you cannot access websites from "the free world" because of these laws, it means that the providers of said websites are threatening your freedom, not providing it.

> China is a suppressive state that strives to control its citizens

China's central government also believes it is protecting its citizens.

> while the EU privacy protection laws are put in place to protect citizens

The fact that they CAN exert so much power on information access in the name of "protection" is a bad precedent, and opens the door to future, less-benevolent authoritarian leadership being formed.

(Even if you think they are protecting their citizens now, I actually disagree; blocking access to AI isn't protecting its citizens, it's handicapping them in the face of a rapidly-advancing world economy.)

Re: Meta Llama 3

#172
post #4

They've got a console for it as well, https://www.meta.ai/ And announcing a lot of integration across the Meta product suite, https://about.fb.com/news/2024/04/meta-ai-assistant-built-wi... Neglected to include comparisons against GPT-4-Turbo or Claude Opus, so I guess it's far from being a frontier model. We'll see how it fares in the LLM Arena.

> Neglected to include comparisons against GPT-4-Turbo or Claude Opus, so I guess it's far from being a frontier model Yeah, almost like comparing a 70b model with a 1.8 trillion parameter model doesn't make any sense when you have a 400b model pending release.

(You can't compare parameter count with a mixture of experts model, which is what the 1.8T rumor says that GPT-4 is.)

Re: Meta Llama 3

#173
post #154

Earlier quoted context omitted.

Wild considering, GPT-4 is 1.8T.

Once benchmarks exist for a while, they become meaningless - even if it's not specifically training on the test set, actions (what used to be called "graduate student descent") end up optimizing new models towards overfitting on benchmark tasks.

Also, the technological leader focuses less on the benchmarks

Re: Meta Llama 3

#174

Would love to experiment with this for work, but the following clause in the license (notably absent in the Llama 2 license) would make this really hard: > i. If you distribute or make available the Llama Materials (or any derivative works thereof), or a product or service that uses any of them, including another AI model, you shall (A) provide a copy of this Agreement with any such Llama Materials; and (B) prominent…

This is the mildest possible clause they could have included short of making the whole thing public domain. Heck the MIT license has similar requirements ("The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.")

Re: Meta Llama 3

#176
post #4

They've got a console for it as well, https://www.meta.ai/ And announcing a lot of integration across the Meta product suite, https://about.fb.com/news/2024/04/meta-ai-assistant-built-wi... Neglected to include comparisons against GPT-4-Turbo or Claude Opus, so I guess it's far from being a frontier model. We'll see how it fares in the LLM Arena.

And they even allow you to use it without logging in. Didnt expect that from Meta.

Doesn't work for me, I'm in EU.

Re: Meta Llama 3

#177

Earlier quoted context omitted.

Where did you find this number? Not doubting it, just want to get a better idea of how precise the estimate may be.

It's a very plausible rumor, but it is misleading in this context, because the rumor also states that it's a mixture of experts model with 8 experts, suggesting that most (perhaps as many as 7/8) of those weights are unused by any particular inference pass. That might suggest that GPT-4 should be thought of as something like a 250B model. But there's also some selection for the remaining 1/8 of weights that are used…

But from an output quality standpoint the total parameter count still seems more relevant. For example 8x7B Mixtral only executes 13B parameters per token, but it behaves comparable to 34B and 70B models, which tracks with its total size of ~45B parameters. You get some of the training and inference advantages of a 13B model, with the strength of a 45B model.

Similarly, if GPT-4 is really 1.8T you would expect it to produce output of similar quality to a comparable 1.8T model without MoE architecture.

Re: Meta Llama 3

#178

Earlier quoted context omitted.

Where did you find this number? Not doubting it, just want to get a better idea of how precise the estimate may be.

It's a very plausible rumor, but it is misleading in this context, because the rumor also states that it's a mixture of experts model with 8 experts, suggesting that most (perhaps as many as 7/8) of those weights are unused by any particular inference pass. That might suggest that GPT-4 should be thought of as something like a 250B model. But there's also some selection for the remaining 1/8 of weights that are used…

I think its almost certainly using at least two experts per token. It helps a lot during training to have two experts to contrast when putting losses on the expert router.

Re: Meta Llama 3

#179
post #168
post #161

Earlier quoted context omitted.

[flagged]

What a silly, provocative comparison. China is a suppressive state that strives to control its citizens while the EU privacy protection laws are put in place to protect citizens. If you cannot access websites from "the free world" because of these laws, it means that the providers of said websites are threatening your freedom, not providing it.

[flagged]

Re: Meta Llama 3

#180
What sort of hardware is needed to run either of these models in a usable fashion? I suppose the bigger 70B model is completely unusable for regular mortals...
Post reply on HN