Live data from Hacker News

Meta Llama 3

llama.meta.com

291–300 of 965 posts

Re: Meta Llama 3

#291

Earlier quoted context omitted.

I agree that it is the best AI podcast. I do have a few gripes though, which might just be from personal preference. A lot of the time the language used by both the host and the guests is unnecessarily obtuse. Also the host is biased towards being optimistic about LLMs leading to AGI, and so he doesn't probe guests deep enough about that, more than just asking something along the lines of "Do you think next token pre…

There's a difference to being a good chatshow/podcast host and a journalist holding someone's feet to the fire! Dwarkesh is excellent at what he does - lots of research beforehand (which is how he lands these great guests), but then lets the guest do most of the talking, and encourages them to expand on what they are saying. It you are critisizing the guest or giving them too much push back, then they are going to cl…

I haven't listened to Dwarkesh, but I take the complaint to mean that he doesn't probe his guests in interesting ways, not so much that he doesn't criticize his guests. If you aren't guiding the conversation into interesting corners then that seems like a problem.

Re: Meta Llama 3

#292
post #58

Earlier quoted context omitted.

I can't express how good Dwarkesh's podcast is in general.

Lex walked so that Dwarkesh could run. He runs the best AI podcast around right now, by a long shot.

indeed my thoughts, especially with first Dario Amodei's interview. He was able to ask all the right questions and discussion was super fruitful.

Re: Meta Llama 3

#293

How does it make monetary sense to release open source models? AFAIK it's very expensive to train them. Do Meta/Mistral have any plans to monetize them?

they are rolling them into the platform, they will obviously boost their ad sales

Re: Meta Llama 3

#294

Earlier quoted context omitted.

Lex walked so that Dwarkesh could run. He runs the best AI podcast around right now, by a long shot.

I don't know Dwarkesh but I despise Lex Fridman. I don't know how a man that lacks the barest modicum of charisma has propelled himself to helming a high-profile, successful podcast. It's not like he tends to express interesting or original thoughts to make up for his paucity of presence. It's bizarre. Maybe I'll check out Dwarkesh, but even seeing him mentioned him in the same breath as Fridman gives me pause ...

I agree with you so much, but he has a solid programmatic approach, where some of the guests uncover. Maybe that's the whole role of an interviewer.

Re: Meta Llama 3

#295
post #279

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

Why is Meta doing it though? This is an astronomical investment. What do they gain from it?

If they start selling ai in their platform, it's a really good option, as people know they can run it somewhere else if they had to (for any reason, e.g: you could make a poc with their platform but then because of regulations you need to self host, can you do that with other offers?)

Re: Meta Llama 3

#296
post #171
post #168

Earlier quoted context omitted.

What a silly, provocative comparison. China is a suppressive state that strives to control its citizens while the EU privacy protection laws are put in place to protect citizens. If you cannot access websites from "the free world" because of these laws, it means that the providers of said websites are threatening your freedom, not providing it.

> China is a suppressive state that strives to control its citizens China's central government also believes it is protecting its citizens. > while the EU privacy protection laws are put in place to protect citizens The fact that they CAN exert so much power on information access in the name of "protection" is a bad precedent, and opens the door to future, less-benevolent authoritarian leadership being formed. (Even…

> The fact that they CAN exert so much power on information access in

They don't have any power on information access. They just require their citizen can decide what you do with it. There is no central system where information is stored that can be used in future by authoritarian leadership. But the information stored about American by American companies can be use in such a way if there one day an authoritarian leadership in America.

Re: Meta Llama 3

#297

If any one is interesting in seeing how 400B model compares with other opensource models, here is a useful chart: https://x.com/natolambert/status/1780993655274414123

Fun fact, it's impossible to 100% the MMLU because 2-3% of it has wrong answers.

Re: Meta Llama 3

#298
post #240

I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…

You can see from Zuck's interviews that he is still an engineer at heart. Every other big tech company has lost that kind of leadership.

[flagged]

Re: Meta Llama 3

#299

Earlier quoted context omitted.

What is "source" regarding an LLM? Public training data and initial parameters?

See this discussion and blog post about a model called OLMo from AI2 ( https://news.ycombinator.com/item?id=39974374 ). They try to be more truly open, although here are nuances even with them that make it not fully open. Just like with open source software, an open source model should provide everything you need to reproduce the final output, and with transparency. That means you need the training source code, the d…

what are you thoughts on projects like these: https://www.llm360.ai/

seems like they make everything available.

Re: Meta Llama 3

#300

Interesting to see that their model comparisons don’t include OpenAI models.

Maybe not the reason, but claude sonnet obliterates gpt3.5 and there isn't a direct llama competitor to gpt4.

The 400B model seems to be a competitor, maybe not in parameter count, but benchmark-wise it seems to be similar.
Post reply on HN