Earlier quoted context omitted.
Well, they're not AI companies, necessarily, or at least not only AI companies, but the big hardware firms tend to have engineers at the helm. That includes Nvidia, AMD, and Intel. (Counterpoint: Apple)
Counter counter point: apples hardware division has been doing great work in the last 5 years, it’s their software that seems to have gone off the rails (in my opinion).
Meta Llama 3
421–430 of 965 posts
Re: Meta Llama 3
#422Earlier quoted context omitted.
Llama is open weight, not open source. They don’t release all the things you need to reproduce their weights.
Not really that either, if we assume that “open weight” means something similar to the standard meaning of “open source”—section 2 of the license discriminates against some users, and the entirety of the AUP against some uses, in contravention of FSD #0 (“The freedom to run the program as you wish, for any purpose”) as well as DFSG #5&6 = OSD #5&6 (“No Discrimination Against Persons or Groups” and “... Fields of Ende…
Re: Meta Llama 3
#423Earlier quoted context omitted.
I can't express how good Dwarkesh's podcast is in general.
Lex walked so that Dwarkesh could run. He runs the best AI podcast around right now, by a long shot.
There is no real commentary to pull from his interviews, at best you get some interesting stories but not the truth.
Re: Meta Llama 3
#424Earlier quoted context omitted.
> his contributions to ... raising salaries It's fun to be able to retire early or whatever, but driving software engineer salaries out of reach of otherwise profitable, sustainable businesses is not a good thing. That just concentrates the industry in fewer hands and makes it more dependent on fickle cash sources (investors, market expansion) often disconnected from the actual software being produced by their teams.…
> It's fun to be able to retire early or whatever, but driving software engineer salaries out of reach of otherwise profitable, sustainable businesses is not a good thing. That argument could apply to anyone who pays anyone well. Driving up market pay for workers via competition for their labour is exactly how we get progress for workers. (And by 'treat well', I mean the whole package. Fortunately, or unfortunately,…
There's a difference between "paying higher salaries in fair competition for talents" and "buying people to let them rot to make sure they don't work for competition".
It's the same as "lowering prices to the benefit of consumer" vs "price dumping to become a monopoly".
Facebook never did it at scale though. Google did.
Re: Meta Llama 3
#425I just want to express how grateful I am that Zuck and Yann and the rest of the Meta team have adopted an open approach and are sharing the model weights, the tokenizer, information about the training data, etc. They, more than anyone else, are responsible for the explosion of open research and improvement that has happened with things like llama.cpp that now allow you to run quite decent models locally on consumer h…
You can see from Zuck's interviews that he is still an engineer at heart. Every other big tech company has lost that kind of leadership.
1. It attracts the world's best academic talent, who deeply want their work shared. AI experts can join any company, so ones which commit to open AI have a huge advantage.
2. Having armies of SWEs contributing millions of free labor hours to test/fix/improve/expand your stuff is incredible.
3. The industry standardizes around their tech, driving down costs and dramatically improving compatibility/extensibility.
4. It creates immense goodwill with basically everyone.
5. Having open AI doesn't hurt their core business. If you're an AI company, giving away your only product isn't tenable (so far).
If Meta's 405B model surpasses GPT-4 and Claude Opus as they expect, they release it for free, and (predictably) nothing awful happens -- just incredible unlocks for regular people like Llama 2 -- it'll make much of the industry look like complete clowns. Hiding their models with some pretext about safety, the alarmist alignment rhetoric, will crumble. Like...no, you zealously guard your models because you want to make money, and that's fine. But using some holier-than-thou "it's for your own good" public gaslighting is wildly inappropriate, paternalistic, and condescending.
The 405B model will be an enormous middle finger to companies who literally won't even tell you how big their models are (because "safety", I guess). Here's a model better than all of yours, it's open for everyone to benefit from, and it didn't end the world. So go &%$# yourselves.
Re: Meta Llama 3
#426Re: Meta Llama 3
#427Where are f32 and f16 used? I see a lot of `.float()' and `.type_as()' in the model file, and nothing explicit about f16. Are the weights and all the activations in f32?
Re: Meta Llama 3
#428Re: Meta Llama 3
#429Earlier quoted context omitted.
His engineering mindset made him blind to the fact the metaverse was a product that nobody wanted or needed. In one of the Fridman interviews, he goes on and on about all the cool technical challenges involved in making the metaverse work. But when Fridman asked him what he likes to do in his spare time, it was all things that you could precisely not do in the metaverse. It was baffling to me that he failed to connec…
I don't think that was the issue. VRChat was basically the same idea but done in a more appealing way and it was (still is) wildly popular.
Re: Meta Llama 3
#430Earlier quoted context omitted.
Depends on your size threshhold. For anything beyond 100 bn in market cap certainly. There is some relatively large companies with a similar flair though, like Cohere and obviously Mistral.
Well, they're not AI companies, necessarily, or at least not only AI companies, but the big hardware firms tend to have engineers at the helm. That includes Nvidia, AMD, and Intel. (Counterpoint: Apple)