Live data from Hacker News

Meta Llama 3

llama.meta.com

471–480 of 965 posts

Re: Meta Llama 3

#471
post #279

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

Why is Meta doing it though? This is an astronomical investment. What do they gain from it?

Besides everything said here in comments, Zuck would be actively looking to own the next platform (after desktop/laptop and mobile), and everyone's trying to figure what that would be.

He knows well that if competitors have a cash cow, they have $$ to throw at hundreds of things. By releasing open-source, he is winning credibility, establishing Meta as the most used LLM, and finally weakening the competition from throwing money on the future initiatives.

Re: Meta Llama 3

#472

Earlier quoted context omitted.

> his contributions to ... raising salaries It's fun to be able to retire early or whatever, but driving software engineer salaries out of reach of otherwise profitable, sustainable businesses is not a good thing. That just concentrates the industry in fewer hands and makes it more dependent on fickle cash sources (investors, market expansion) often disconnected from the actual software being produced by their teams.…

> but driving software engineer salaries out of reach of otherwise profitable, sustainable businesses is not a good thing. I'm not convinced he's actually done that. Pretty much any 'profitable, sustainable business' can afford software developers. Software developers are paid pretty decently, but (grabbing a couple of lists off of Google) it looks like there's 18 careers more lucrative than it (from a wage perspecti…

Few viable technology businesses and non-technology busiesses with internal software departments were prepared to see their software engineers suddenly suddenly expect doctor or lawyer pay and can't effectively accomodate the change.

They were largely left to rely on loyalty and other kinds of fragile non-monetary factors to preserve their existing talent and institutuonal knowledge and otherwise scavenge for scraps when making new hires.

For those companies outside the specific Silicon Valley money circle, it was an extremely disruptive change and recovery basically requires that salaries normalize to some significant degree. In most cases, engineers provide quite a lot of value but not nearly so much value as FAANG and SV speculators could build into their market-shaping offers.

It's not a healthy situation for the industry or (if you're wary of centralization/monopolization) society as a whole.

Re: Meta Llama 3

#473

Very strong results for their size on my NYT Connections benchmark. Llama 3 Instruct 70B better than new commercial models Gemini Pro 1.5 and Mistral Large and not far away from Clause 3 Opus and GPT-4. Llama 3 Instruct 8B better than larger open weights models like Mixtral-8x22B. Full list: https://twitter.com/LechMazur/status/1781049810428088465/pho...

Cool, I enjoy doing Connections! Do you have a blog post or github code available? Or do you stick to only xeets?

Re: Meta Llama 3

#474

I was curious how the numbers compare to GPT-4 in the paid ChatGPT Plus, since they don't compare directly themselves. Llama 3 8B Llama 3 70B GPT-4 MMLU 68.4 82.0 86.5 GPQA 34.2 39.5 49.1 MATH 30.0 50.4 72.2 HumanEval 62.2 81.7 87.6 DROP 58.4 79.7 85.4 Note that the free version of ChatGPT that most people use is based on GPT-3.5 which is much worse than GPT-4. I haven't found comprehensive eval numbers for the lates…

But I'm waiting for the finetunedz/merged models. Many devs produced great models based on Llama 2, that outperformed the vanilla one, so I expect similar treatment for the new version. Exciting nonetheless!

Re: Meta Llama 3

#475

Earlier quoted context omitted.

The parameters and the license. Mistral uses Apache 2.0, a neatly permissive open source license. As such, it's an open source model. Models are similar to code you might run on a compiled vm or native operating system. Llama.cpp is to a model as Python is to a python script. The license lays out the rights and responsibilities of the users of the software, or the model, in this case. The training data, process, pipe…

Mistral is not “open source” either since we cannot reproduce it (the training data is not published). Both are open weight models, and they are both released under a license whose legal basis is unclear: it's not actually clear if they own any intellectual property over the model at all. Of course they claim such IP, but no court has ruled on this yet AFAIK and legislators could also enact laws that make these publi…

Is “reproducibility” actually the right term here?

It’s a bit like arguing that Linux is not open source because you don’t have every email Linus and the maintainers ever received. Or that you don’t know what lectures Linus attended or what books he’s read.

The weights “are the thing” in the same sense that the “code is the thing”. You can modify open code and recompile it. You can similarly modify weights with fine tuning or even architectural changes. You don’t need to go “back to the beginning” in the same sense that Linux would continue to be open source even without the Git history and the LKM mailing list.

Re: Meta Llama 3

#477

Earlier quoted context omitted.

The parameters and the license. Mistral uses Apache 2.0, a neatly permissive open source license. As such, it's an open source model. Models are similar to code you might run on a compiled vm or native operating system. Llama.cpp is to a model as Python is to a python script. The license lays out the rights and responsibilities of the users of the software, or the model, in this case. The training data, process, pipe…

Mistral is not “open source” either since we cannot reproduce it (the training data is not published). Both are open weight models, and they are both released under a license whose legal basis is unclear: it's not actually clear if they own any intellectual property over the model at all. Of course they claim such IP, but no court has ruled on this yet AFAIK and legislators could also enact laws that make these publi…

I have a hard time about the "cannot reproduce" categorization.

There are places (e.g. in the Linux kernel? AMD drivers?) where lots of generated code is pushed and (apart from the rants of huge unwieldy commits and complaints that it would be better engineering-wise to get their hands on the code generator, it seems no one is saying the AMD drivers aren't GPL compliant or OSI-compliant?

There are probably lots of OSS that is filled with constants and code they probably couldn't rederive easily, and we still call them OSS?

Re: Meta Llama 3

#478
what's the state of the art in quantization methods these days that one might apply to a model like LLama 3? Any particular literature to read? Of course priorities differ across methods. Rather than saving space or speeding up calculations, I'm simply interested in static quantization where integer weights multiply integer activations (like 8-bit integers). (as for motivation, such quantization enables proving correct execution of inference in sublinear time, at least asymptotically. i'm talking of ZK tech)

Re: Meta Llama 3

#479
It’s amazing seeing everyone collectively trust every company over and over again only to get burned over and over again. I can’t wait for Meta to suddenly lock down newer versions after they’ve received enough help from everyone else, just so that developers can go omg who could’ve ever predicted this?

Re: Meta Llama 3

#480

Earlier quoted context omitted.

Yes. And, could potentially diminish OpenAI/MS. Once everyone can do it, then OpenAI value would evaporate.

Once every human has access to cutting edge AI, that ceases to be a differentiating factor, so the human talent will again be the determining factor.

And the content industry will grow ever more addictive and profitable, with content curated and customized specifically for your psyche. The very industry Meta happens to be the one to benefit from its growth most among all tech giants.
Post reply on HN