Live data from Hacker News

Meta Llama 3

llama.meta.com

531–540 of 965 posts

Re: Meta Llama 3

#531
post #256
post #173

Earlier quoted context omitted.

Also, the technological leader focuses less on the benchmarks

Interesting claim, is there data to back this up? My impression is that Intel and NVIDIA have always gamed the benchmarks.

NVIDIA needs T models not B models to keep the share price up.

Re: Meta Llama 3

#532
post #240

Earlier quoted context omitted.

You can see from Zuck's interviews that he is still an engineer at heart. Every other big tech company has lost that kind of leadership.

NVidia, AMD, Microsoft?

Nvidia, maybe. Microsoft, definitely not. Nadella is a successful CEO but is as corporate as they come.

Re: Meta Llama 3

#533
post #171
post #168

Earlier quoted context omitted.

What a silly, provocative comparison. China is a suppressive state that strives to control its citizens while the EU privacy protection laws are put in place to protect citizens. If you cannot access websites from "the free world" because of these laws, it means that the providers of said websites are threatening your freedom, not providing it.

> China is a suppressive state that strives to control its citizens China's central government also believes it is protecting its citizens. > while the EU privacy protection laws are put in place to protect citizens The fact that they CAN exert so much power on information access in the name of "protection" is a bad precedent, and opens the door to future, less-benevolent authoritarian leadership being formed. (Even…

>China's central government also believes it is protecting its citizens.

Anyone who's taking a course in epistemology can tell you that there's more to assessing veracity of a belief than noting its equivalence to other beliefs. There can be symmetry in psychology without symmetry in underlying facts. So noting an equivalence of belief is not enough to establish an equivalence in fact.

I'm not even saying I'm for or against the EU's choices but I think the purpose of analogies to China is kind of rhetorical purpose of warning or a comparison intended to reflect negatively on the EU. I find it hard to imagine one would make a straight faced case that they are in fact equivalent in scope or scale or ambition or equivalent and their idea of the relation of their mission to their values for core liberties.

I think the difference is here are clear enough that reasonable people should be able to make the case against AI regulation without losing grasp of the distinction between European and Chinese regulatory frameworks.

Re: Meta Llama 3

#534

Quick thoughts - Major arch changes are not that major, mostly GQA and tokenizer improvements. Tokenizer improvement is a under-explored domain IMO. 15T tokens is a ton! 400B model performance looks great, can’t wait for that to be released. Might be time to invest in a Mac studio! OpenAI probably needs to release GPT-5 soon to convince people they are still staying ahead.

> Might be time to invest in a Mac studio! The highest end Mac Studio with 196GB of ram won't even be enough to run a Q4 quant of the 400B+ (don't forget the +) model. At this point, one have to consider an Epyc for CPU inference or costlier gpu solutions like the "popular" 8xA100 80GB... An if it's a dense model like the other llamas, it will be pretty slow..

Just FYI on the podcast video Zuck seems to let it slip that the exact number is 405B. (2-3mins in)

Re: Meta Llama 3

#535
post #378
post #349

Earlier quoted context omitted.

he said "Q4" - meaning 4-bit weights.

Ok but at 16-bit it would be 800GB+, right? Not 512.

Divide not multiply. If a size is estimated in 8-bit, reducing to 4-bit halves the size (and entropy of each value). Difference between INT_MAX and SHORT_MAX (assuming you have such defs).

I could be wrong too but that’s my understanding. Like float vs half-float.

Re: Meta Llama 3

#536
post #93

Earlier quoted context omitted.

They didn't compare against the best models because they were trying to do "in class" comparisons, and the 70B model is in the same class as Sonnet (which they do compare against) and GPT3.5 (which is much worse than sonnet). If they're beating sonnet that means they're going to be within stabbing distance of opus and gpt4 for most tasks, with the only major difference probably arising in extremely difficult reasonin…

ML Twitter was saying that they're working on a 400B parameter version?

Meta themselves are saying that: https://ai.meta.com/blog/meta-llama-3/

Re: Meta Llama 3

#537

Earlier quoted context omitted.

Also, being open source adds phenomenal value for Meta: 1. It attracts the world's best academic talent, who deeply want their work shared. AI experts can join any company, so ones which commit to open AI have a huge advantage. 2. Having armies of SWEs contributing millions of free labor hours to test/fix/improve/expand your stuff is incredible. 3. The industry standardizes around their tech, driving down costs and d…

OpenAI engineers don't work for free. Facebook subsidizes their engineers because they have $20B. OpenAI doesn't have that luxury.

Sucks to work in a non-profit, right? Oh wait... }:^). Those assholes are lobbying to block public llm, 0 sympathy.

Re: Meta Llama 3

#538

From the article >We made several new observations on scaling behavior during the development of Llama 3. For example, while the Chinchilla-optimal amount of training compute for an 8B parameter model corresponds to ~200B tokens, we found that model performance continues to improve even after the model is trained on two orders of magnitude more data. Both our 8B and 70B parameter models continued to improve log-linea…

Yes. Llama 3 8B outperforms Llama 2 70B (in the instruct-tuned variants). "Chinchilla-optimal" is about choosing model size and/or dataset size to maximize the accuracy of your model under a fixed training budget (fixed number of floating point operations). For a given dataset size it will tell you the model size to use, and vice versa, again under the assumption of a fixed training budget. However, what people have…

Somewhere I read that the 8B llama2 model could be undertrained by 100-1000x. So is it possible to train a model with 8B/100 = 80M parameters to perform as good as the llama2 8B model, given enough training time and training tokens?

Re: Meta Llama 3

#539
post #8

> We’re rolling out Meta AI in English in more than a dozen countries outside of the US. Now, people will have access to Meta AI in Australia, Canada, Ghana, Jamaica, Malawi, New Zealand, Nigeria, Pakistan, Singapore, South Africa, Uganda, Zambia and Zimbabwe — and we’re just getting started.

ie America + a selection of countries that mostly haven’t got their shit together yet on dealing with the threat of unregulated AI.
Post reply on HN