Live data from Hacker News

Meta Llama 3

llama.meta.com

321–330 of 965 posts

Re: Meta Llama 3

#321
post #112

Earlier quoted context omitted.

The bottom of https://ai.meta.com/blog/meta-llama-3/ has in-progress results for the 400B model as well. Looks like it's not quite there yet. Llama 3 400B Base / Instruct MMLU 84.8 86.1 GPQA - 48.0 MATH - 57.8 HumanEval - 84.1 DROP 83.5 -

For the still training 400B: Llama 3 GPT 4(Published) BBH 85.3 83.1 MMLU 86.1 86.4 DROP 83.5 80.9 GSM8K 94.1 92.0 MATH 57.8 52.9 HumEv 84.1 74.4 Although it should be noted that the API numbers were generally better than published numbers for GPT4. [1]: https://deepmind.google/technologies/gemini/

Which specific GPT-4 model is this? gpt-4-0613? gpt-4-0125-preview?

Re: Meta Llama 3

#322

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

They're sharing it for a reason. That reason is to disarm their opponents.

Re: Meta Llama 3

#323

I was curious how the numbers compare to GPT-4 in the paid ChatGPT Plus, since they don't compare directly themselves. Llama 3 8B Llama 3 70B GPT-4 MMLU 68.4 82.0 86.5 GPQA 34.2 39.5 49.1 MATH 30.0 50.4 72.2 HumanEval 62.2 81.7 87.6 DROP 58.4 79.7 85.4 Note that the free version of ChatGPT that most people use is based on GPT-3.5 which is much worse than GPT-4. I haven't found comprehensive eval numbers for the lates…

Via Microsoft Copilot (and perhaps Bing?) you can get access to GPT-4 for free.

Re: Meta Llama 3

#324

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

Good thing that he's only 39 years old and seems more energetic than ever to run his company. Having a passionate founder is, imo, a big advantage for Meta compared to other big tech companies.

Love how everyone is romanticizing his engineering mindset. But have we already forgotten that he was even more passionate about the metaverse which, as far as I can tell, was a 50B failure?

Re: Meta Llama 3

#326

Quick thoughts - Major arch changes are not that major, mostly GQA and tokenizer improvements. Tokenizer improvement is a under-explored domain IMO. 15T tokens is a ton! 400B model performance looks great, can’t wait for that to be released. Might be time to invest in a Mac studio! OpenAI probably needs to release GPT-5 soon to convince people they are still staying ahead.

> Might be time to invest in a Mac studio!

it's wild isn't it

for so long a few years old macbook is fine for everything, in desperation Apple waste their time with VR goggles in search of a use-case... then suddenly ChatGPT etc comes along and despite relatively weak GPU Apple accidentally have stuff worth upgrading to

imagine when they eventually take the goggles off and start facing in the right direction...

Re: Meta Llama 3

#327
They've added a big, colorful, ugly button to my WhatsApp now. At the moment the button is covering the date information of my last chat with my Mom. It's revolting.

Re: Meta Llama 3

#328
post #279

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

Why is Meta doing it though? This is an astronomical investment. What do they gain from it?

He went into the details of how he thinks about open sourcing weights for Llama responding to a question from an analyst in one of the earnings call last year after Llama release. I had made a post on Reddit with some details.

https://www.reddit.com/r/MachineLearning/s/GK57eB2qiz

Some noteworthy quotes that signal the thought process at Meta FAIR and more broadly

* We’re just playing a different game on the infrastructure than companies like Google or Microsoft or Amazon

* We would aspire to and hope to make even more open than that. So, we’ll need to figure out a way to do that.

* ...lead us to do more work in terms of open sourcing, some of the lower level models and tools

* Open sourcing low level tools make the way we run all this infrastructure more efficient over time.

* On PyTorch: It’s generally been very valuable for us to provide that because now all of the best developers across the industry are using tools that we’re also using internally.

* I would expect us to be pushing and helping to build out an open ecosystem.

Re: Meta Llama 3

#329

Earlier quoted context omitted.

Lex walked so that Dwarkesh could run. He runs the best AI podcast around right now, by a long shot.

I don't know Dwarkesh but I despise Lex Fridman. I don't know how a man that lacks the barest modicum of charisma has propelled himself to helming a high-profile, successful podcast. It's not like he tends to express interesting or original thoughts to make up for his paucity of presence. It's bizarre. Maybe I'll check out Dwarkesh, but even seeing him mentioned him in the same breath as Fridman gives me pause ...

The question you should ask is: why are high-profile guests willing to talk to Lex Fridman but not others?

The short answer, imho, is trust. No one gets turned into an embarrassing soundbite talking to Lex. He doesn't try to ask gotcha questions for clickbait articles. Generally speaking "the press" are not your friend and they will twist your words. You have to walk on egg shells.

Lex doesn't need to express original ideas. He needs to get his guests to open up and share their unique perspectives and thoughts. He's been extremely successful in this.

An alternative question is why hasn't someone more charismatic taken off in this space? I'm not sure! Who knows, there might be some lizard brain secret sauce behind the "flat" podcast host.

Re: Meta Llama 3

#330

The instant generation of pictures as you type in meta.ai is really impressive!

It is. But I noticed something weird. If your prompt is “A cartoon of XYZ” and press enter the preview will be a cartoon but the other images will be weird realistic ones.

The preview is using a different faster model so you're not going to get the exact same styles of responses from the larger slower one. If you have ideas on how to make the user experience better based on those constraints please let us know!
Post reply on HN