Live data from Hacker News

Meta Llama 3

llama.meta.com

341–350 of 965 posts

Re: Meta Llama 3

#341
post #317

Earlier quoted context omitted.

Meta is an advertising company that is primarily driven by user generated content. If they can empower more people to create more content more quickly, they make more money. Particularly the metaverse, if they ever get there, because making content for 3d VR is very resource intensive. Making AI as open as possible so more people can use it accelerates the rate of content creation

You could say the same about Google, couldn't you?

Yea probably, but I don't think Google as a company is trying to do anything open regarding AI other than raw research papers

Also google makes most of their money off search, which is more business driven advertising vs showing ads in between user generated content bites

Re: Meta Llama 3

#342

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

It does seem uncharacteristic. Wonder how much of the hate Zuck gets is people that just don't like Facebook, but as person/engineer, his heart is in the right place? It is hard to accept this at face value and not think there is some giant corporate hidden agenda.

Re: Meta Llama 3

#343

Earlier quoted context omitted.

Good thing that he's only 39 years old and seems more energetic than ever to run his company. Having a passionate founder is, imo, a big advantage for Meta compared to other big tech companies.

Love how everyone is romanticizing his engineering mindset. But have we already forgotten that he was even more passionate about the metaverse which, as far as I can tell, was a 50B failure?

Think of it as a 50B spending spree where he gave that to VR tech out of enthusiasm. Even I, with the cold dark heart that I have, has to admit he's a geek hero with his open source attitude.

Re: Meta Llama 3

#344

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

I kind of wonder. Does what they do counter the growth of Google?

I remember reading years ago that page/brin wanted to build an AI.

This was long before the AI boom, when saying something like that was just weird (like musk saying he wanted to die on mars weird)

Re: Meta Llama 3

#345
post #154

Earlier quoted context omitted.

Wild considering, GPT-4 is 1.8T.

Once benchmarks exist for a while, they become meaningless - even if it's not specifically training on the test set, actions (what used to be called "graduate student descent") end up optimizing new models towards overfitting on benchmark tasks.

Even random seed could cause bad big shift in human eval performance if you know you know. It is perfectly illegal to choose one ckpt that looks best on those benchmarks and move along

HumanEval is meaningless regardless, those 164 problems have been overfit to the tea.

Hook this up to LLM arena we will get a better picture regarding how powerful they really are

Re: Meta Llama 3

#346
post #16

Zuck has an interview out for it as well, https://twitter.com/dwarkesh_sp/status/1780990840179187715

Seems like a year or two of MMA has done way more for his charisma than whatever media training he's done over the years. He's a lot more natural in interviews now.

Intense exercise, especially a competetive sport where you train with other people tends to do this

Re: Meta Llama 3

#347

Earlier quoted context omitted.

Good thing that he's only 39 years old and seems more energetic than ever to run his company. Having a passionate founder is, imo, a big advantage for Meta compared to other big tech companies.

Love how everyone is romanticizing his engineering mindset. But have we already forgotten that he was even more passionate about the metaverse which, as far as I can tell, was a 50B failure?

It isn't necessarily a failure "yet". Don't think anybody is saying VR/AR isn't a huge future product, just that current tech is not quite there. We'll see if Apple can do better, they both made tradeoffs.

It is still possible that VR and Generative AI can join in some synergy.

Re: Meta Llama 3

#348
post #277

Earlier quoted context omitted.

My personal theory is that this is all because Zuckerberg has a rivalry with Elon Musk, who is an AI decelerationist (well, when it's convenient for him) and appears to believe in keeping AI in the control of the few. There was a spat between them a few years ago on Twitter where Musk said Zuckerberg had limited understanding of AI tech, after Zuckerberg called out AI doomerism as stupid.

It's a silly but spooky thought that this or similar interactions may have been the butterfly effect that drove at least one of them to take their company in a drastically different direction.

There's probably all sorts of things that happen for reasons we'll never know. These are both immensely powerful men driven by ego and the idea of leaving a legacy. It's not unreasonable to think one of them might throw around a few billion just to spite the other.

Re: Meta Llama 3

#349
post #282

Earlier quoted context omitted.

My guess is around 256GiB but it depends on what level of quantization you are okay with. At full 16bit it will be massive, near 512GiB. I figure we will see some Q4's that can probably fit on 4 4090s with CPU offloading.

With 400 billion parameters and 8 bits per parameter, wouldn't it be ~400 GB? Plus context size which could be quite large.

he said "Q4" - meaning 4-bit weights.

Re: Meta Llama 3

#350
post #279

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

Why is Meta doing it though? This is an astronomical investment. What do they gain from it?

It’s a shame it can’t just be giving back to the community and not questioned.

Why is selfishness from companies who’ve benefited from social resources not a surprising event vs the norm.

Post reply on HN