Live data from Hacker News

Meta Llama 3

llama.meta.com

561–570 of 965 posts

Re: Meta Llama 3

#561
post #279

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

Why is Meta doing it though? This is an astronomical investment. What do they gain from it?

I think you really have to understand Zuckerberg's "origin story" to understand why he is doing this. He created a thing called Facebook that was wildly successful. Built it with his own two hands. We all know this.

But what is less understood is that from his point of view, Facebook went through a near death experience when mobile happened. Apple and Google nearly "stole" it from him by putting strict controls around the next platform that happened, mobile. He lives every day even still knowing Apple or Google could simply turn off his apps and the whole dream would come to an end.

So what do you do in that situation? You swear - never again. When the next revolution happens, I'm going to be there, owning it from the ground up myself. But more than that, he wants to fundamentally shift the world back to the premise that made him successful in the first place - open platforms. He thinks that when everyone is competing on a level playing field he'll win. He thinks he is at least as smart and as good as everyone else. The biggest threat to him is not that someone else is better, it's that the playing field is made arbitrarily uneven.

Of course, this is all either conjecture or pieced together from scraps of observations over time. But it is very consistent over many decisions and interactions he has made over many years and many different domains.

Re: Meta Llama 3

#562
post #507

Earlier quoted context omitted.

Love how everyone is romanticizing his engineering mindset. But have we already forgotten that he was even more passionate about the metaverse which, as far as I can tell, was a 50B failure?

Having a nerdy vision of the future and spending tens of billions of dollars to try and make it a reality while shareholders and bean counters crucify you for it is the most engineer thing imaginable. What other CEO out there is taking such risks?

Bill Gates when he was at Microsoft.

Tablet PC (first iteration was in the early 90s!), Pocket PC, WebTV and Media Center PC (Microsoft first tried Smart TVs in the late 90s! There wasn't any content to watch and most people didn't have broadband, oops), Xbox, and the numerous PC standards they pushed for (e.g. mandating integrated audio on new PCs), smart watches (SPOT watch, look it up!), and probably a few others I'm forgetting.

You'll notice in most of those categories, they moved too soon and others who came later won the market.

Re: Meta Llama 3

#563
post #552

Earlier quoted context omitted.

Those numbers are for the original GPT-4 (Mar 2023). Current GPT-4-Turbo (Apr 2024) is better: Llama 3 GPT-4 GPT-4-Turbo* (Apr 2024) MMLU 86.1 86.4 86.7 DROP 83.5 80.9 86.0 MATH 57.8 52.9 73.4 HumEv 84.1 74.4 88.2 *using API prompt: https://github.com/openai/simple-evals

I find it somewhat interesting that there is a common perception about GPT-4 at release being actually smart, but that it got gradually nerfed for speed with turbo, which is better tuned but doesn't exhibit intelligence like the original. There were times when I felt that too, but nowadays I predominantly use turbo. It's probably because turbo is faster and cheaper, but in lmsys turbo has 100 elo higher than original…

Given the incremental increase between GPT-4 and its turbo variant, I would weight “vibes” more heavily than this improvement on MMLU. OpenAI isn’t exactly a very honest or transparent company and the metric is imperfect. As a longtime time user of ChatGPT, I observed it got markedly worse at coding after the turbo release, specifically in its refusal to complete code as specified.

Re: Meta Llama 3

#564

Earlier quoted context omitted.

For the still training 400B: Llama 3 GPT 4(Published) BBH 85.3 83.1 MMLU 86.1 86.4 DROP 83.5 80.9 GSM8K 94.1 92.0 MATH 57.8 52.9 HumEv 84.1 74.4 Although it should be noted that the API numbers were generally better than published numbers for GPT4. [1]: https://deepmind.google/technologies/gemini/

Wild! So if this indeed holds up, it looks like OpenAI were about a year ahead when GPT-4 was released, compared to the open source world. However, given the timespan between matching GPT-3.5 (Mixtral perhaps?) and matching GPT-4 has just been a few weeks, I am wondering if the open source models have more momentum. That said, I am very curious what OpenAI has in their labs... Are they actually barely ahead? Or do th…

You've also got to consider that we don't really know where OpenAI are though, what they have released in the past year have been tweaks to GPT4, while I am sure the real work is going into GPT5 or whatever it gets called.

While all the others are catching up and in some cases being slightly better, I wouldn't be surprised to see a rather large leap back into the lead from OpenAI pretty soon and then a scrabble for some time for others to get close again. We will really see who has the momentum soon, when we see OpenAI's next full release.

Re: Meta Llama 3

#565

Earlier quoted context omitted.

Yes. Llama 3 8B outperforms Llama 2 70B (in the instruct-tuned variants). "Chinchilla-optimal" is about choosing model size and/or dataset size to maximize the accuracy of your model under a fixed training budget (fixed number of floating point operations). For a given dataset size it will tell you the model size to use, and vice versa, again under the assumption of a fixed training budget. However, what people have…

Somewhere I read that the 8B llama2 model could be undertrained by 100-1000x. So is it possible to train a model with 8B/100 = 80M parameters to perform as good as the llama2 8B model, given enough training time and training tokens?

It's unclear. It might take a larger dataset than actually exists, or more compute than is practical. Or there may be a limit that we just haven't reached yet; this actually seems quite likely. The scaling "laws" are really more like guidelines and they are likely wrong when extrapolated too far.

Re: Meta Llama 3

#566

Earlier quoted context omitted.

> but driving software engineer salaries out of reach of otherwise profitable, sustainable businesses is not a good thing. I'm not convinced he's actually done that. Pretty much any 'profitable, sustainable business' can afford software developers. Software developers are paid pretty decently, but (grabbing a couple of lists off of Google) it looks like there's 18 careers more lucrative than it (from a wage perspecti…

Few viable technology businesses and non-technology busiesses with internal software departments were prepared to see their software engineers suddenly suddenly expect doctor or lawyer pay and can't effectively accomodate the change. They were largely left to rely on loyalty and other kinds of fragile non-monetary factors to preserve their existing talent and institutuonal knowledge and otherwise scavenge for scraps…

That business can search and find talents globally for fraction of SV salary.

If FAANG company can hire an engineer overseas for 60k$ annually why other cannot?

Re: Meta Llama 3

#567
I asked MetaAI about whether it can take PDF/Image input, and it started hallucinating, giving inconsistent answers. Is it just a bug, or poor prompt engineering, or hallucination?

> Me: can you take pdf?

> Meta AI: Yes, I can take PDFs as input. You can upload a PDF file, and I can extract text from it and respond based on the content. Please upload the PDF file using the "Upload a file" button or copy and paste the text from the PDF into the chat. Remember, I can only access files uploaded or text copied during our current conversation. If you need to refer to a previous conversation, feel free to re-upload the file or re-paste the text. Let me know how I can assist you with the PDF file!

> me: can you take pdf as input?

> Meta AI: I'm unable to directly accept PDF files as input. However, you can copy and paste the text from the PDF file into the chat, and I'll be happy to assist you! Alternatively, if you have a specific question or request related to the content of the PDF file, feel free to ask, and I'll do my best to help!

(the article says meta plans to make it multimodal in the future)

Re: Meta Llama 3

#568

Earlier quoted context omitted.

Somewhere I read that the 8B llama2 model could be undertrained by 100-1000x. So is it possible to train a model with 8B/100 = 80M parameters to perform as good as the llama2 8B model, given enough training time and training tokens?

It's unclear. It might take a larger dataset than actually exists, or more compute than is practical. Or there may be a limit that we just haven't reached yet; this actually seems quite likely. The scaling "laws" are really more like guidelines and they are likely wrong when extrapolated too far.

Thanks!

Re: Meta Llama 3

#569

I asked MetaAI about whether it can take PDF/Image input, and it started hallucinating, giving inconsistent answers. Is it just a bug, or poor prompt engineering, or hallucination? > Me: can you take pdf? > Meta AI: Yes, I can take PDFs as input. You can upload a PDF file, and I can extract text from it and respond based on the content. Please upload the PDF file using the "Upload a file" button or copy and paste the…

[deleted]

Re: Meta Llama 3

#570
Architectural changes between Llama 2 and 3 seem to be minimal. Looking at the 400B model benchmarks and comparing them to GPT-4 only proves that there is no secret sauce. It's all about the dataset and the number of params.
Post reply on HN