Live data from Hacker News

Meta Llama 3

llama.meta.com

311–320 of 965 posts

Re: Meta Llama 3

#311

Earlier quoted context omitted.

Probably the EU laws are getting too draconian. I'm starting to see it a lot.

> the EU laws are getting too draconian You also said that when Meta delayed the Threads release by a few weeks in the EU. I recommend reading the princess on a pea fairytale since you seem to be quite sheltered, using the term draconian as liberally.

>a few weeks

July to December is not "a few weeks"

Re: Meta Llama 3

#312
post #4

They've got a console for it as well, https://www.meta.ai/ And announcing a lot of integration across the Meta product suite, https://about.fb.com/news/2024/04/meta-ai-assistant-built-wi... Neglected to include comparisons against GPT-4-Turbo or Claude Opus, so I guess it's far from being a frontier model. We'll see how it fares in the LLM Arena.

> Meta AI isn't available yet in your country Where is it available? I got this in Norway.

>We’re rolling out Meta AI in English in more than a dozen countries outside of the US. Now, people will have access to Meta AI in Australia, Canada, Ghana, Jamaica, Malawi, New Zealand, Nigeria, Pakistan, Singapore, South Africa, Uganda, Zambia and Zimbabwe — and we’re just getting started.

https://about.fb.com/news/2024/04/meta-ai-assistant-built-wi...

Re: Meta Llama 3

#313
post #279

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

Why is Meta doing it though? This is an astronomical investment. What do they gain from it?

Generative AI is a necessity for the metaverse to take off. Creating metaverse content is too time consuming otherwise. Mark really wants to control a platform so the companies whole strategy seems to be around getting the quest to take off.

Re: Meta Llama 3

#314

From the article >We made several new observations on scaling behavior during the development of Llama 3. For example, while the Chinchilla-optimal amount of training compute for an 8B parameter model corresponds to ~200B tokens, we found that model performance continues to improve even after the model is trained on two orders of magnitude more data. Both our 8B and 70B parameter models continued to improve log-linea…

Yes. Llama 3 8B outperforms Llama 2 70B (in the instruct-tuned variants).

"Chinchilla-optimal" is about choosing model size and/or dataset size to maximize the accuracy of your model under a fixed training budget (fixed number of floating point operations). For a given dataset size it will tell you the model size to use, and vice versa, again under the assumption of a fixed training budget.

However, what people have realized is that inference compute matters at least as much as training compute. You want to optimize training and inference cost together, not in isolation. Training a smaller model means your accuracy will not be as good as it could have been with a larger model using the same training budget, however you'll more than make it up in your inference budget. So in most real world cases it doesn't make sense to be "Chinchilla-optimal".

What Meta is saying here is that there is no accuracy ceiling. You can keep increasing training budget and dataset size to increase accuracy seemingly indefinitely (with diminishing returns). At least as far as they have explored.

Re: Meta Llama 3

#315

Earlier quoted context omitted.

> Meta AI isn't available yet in your country Where is it available? I got this in Norway.

Just use the Replicate demo instead, you can even alter the inference parameters https://llama3.replicate.dev/ Or run a jupyter notebook from Unsloth on Colab https://huggingface.co/unsloth/llama-3-8b-bnb-4bit

This version doesn't have web search and the image creation though.

Re: Meta Llama 3

#316

Earlier quoted context omitted.

The world at large seems to hate Zuck but it’s good to hear from people familiar with software engineering and who understand just how significant his contributions to open source and raising salaries have been through Facebook and now Meta.

> his contributions to ... raising salaries It's fun to be able to retire early or whatever, but driving software engineer salaries out of reach of otherwise profitable, sustainable businesses is not a good thing. That just concentrates the industry in fewer hands and makes it more dependent on fickle cash sources (investors, market expansion) often disconnected from the actual software being produced by their teams.…

> It's fun to be able to retire early or whatever, but driving software engineer salaries out of reach of otherwise profitable, sustainable businesses is not a good thing.

That argument could apply to anyone who pays anyone well.

Driving up market pay for workers via competition for their labour is exactly how we get progress for workers.

(And by 'treat well', I mean the whole package. Fortunately, or unfortunately, that has the side effect of eg paying veterinary nurses peanuts, because there's always people willing to do those kinds of 'cute' jobs.)

> Nor is it great for the yet-to-mature craft that high salaries invited a very large pool of primarly-compensation-motivated people who end up diluting the ability for primarily-craft-motivated people to find and coordinate with each other in pursuit of higher quality work and more robust practices.

Huh, how is that 'dilution' supposed to work?

Well, and at least those 'evil' money grubbers are out of someone else's hair. They don't just get created from thin air. So if those rimarly-compensation-motivated people are now writing software, then at least investment banking and management consulting are free again for the primarily-craft-motivated people to enjoy!

Re: Meta Llama 3

#317
post #279

Earlier quoted context omitted.

Why is Meta doing it though? This is an astronomical investment. What do they gain from it?

Meta is an advertising company that is primarily driven by user generated content. If they can empower more people to create more content more quickly, they make more money. Particularly the metaverse, if they ever get there, because making content for 3d VR is very resource intensive. Making AI as open as possible so more people can use it accelerates the rate of content creation

You could say the same about Google, couldn't you?

Re: Meta Llama 3

#318

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

Meta also spearheaded the open compute project. I originally joined Google because of their commitment to open source and was extremely disappointed when I didn't see that culture continue as we worked on exascale solutions. Glad to see Meta carrying the torch here. Hope it continues.

When did you join Google?

Re: Meta Llama 3

#319

Earlier quoted context omitted.

There's a difference to being a good chatshow/podcast host and a journalist holding someone's feet to the fire! Dwarkesh is excellent at what he does - lots of research beforehand (which is how he lands these great guests), but then lets the guest do most of the talking, and encourages them to expand on what they are saying. It you are critisizing the guest or giving them too much push back, then they are going to cl…

I haven't listened to Dwarkesh, but I take the complaint to mean that he doesn't probe his guests in interesting ways, not so much that he doesn't criticize his guests. If you aren't guiding the conversation into interesting corners then that seems like a problem.

He does a lot of research before his interviews, so comes with a lot of good questions, but then mostly let's the guests talk. He does have some impromptu follow-ups, but mostly tries to come back to his prepared questions.

A couple of his interviews I'd recommend:

- Dario Amodei (Anthropic CEO)

https://www.youtube.com/watch?v=Nlkk3glap_U

- Richard Rhodes (Manhatten project, etc - history of Atom bomb)

https://www.youtube.com/watch?v=tMdMiYsfHKo

Re: Meta Llama 3

#320
post #277
post #267

I'm a big fan of various AI companies taking different approaches. OpenAI keeping it close to their hearts but have great developer apis. Meta and Mistral going open weights + open code. Anthropic and Claude doing their thing. Competition is a beautiful thing. I am half excited and half scared that AGI is our generation's space war. I hope we can solve the big human problems, instead of more scammy ads and videos. So…

My personal theory is that this is all because Zuckerberg has a rivalry with Elon Musk, who is an AI decelerationist (well, when it's convenient for him) and appears to believe in keeping AI in the control of the few. There was a spat between them a few years ago on Twitter where Musk said Zuckerberg had limited understanding of AI tech, after Zuckerberg called out AI doomerism as stupid.

It's a silly but spooky thought that this or similar interactions may have been the butterfly effect that drove at least one of them to take their company in a drastically different direction.
Post reply on HN