Occam’s razor tells me it’s probably because it’s not good. Perhaps running a company like survivor in a pressure cooker is not an effective management strategy.
Seemed to work when it comes to selling ads. I'm thinking training LLMs is harder than anthropic and openai make it look
Meta Keeps Delaying the Release of Its New AI Model to Developers
21–28 of 28 posts
Re: Meta Keeps Delaying the Release of Its New AI Model to Developers
#22I’ve used it at Meta. It’s very bad, if they released it in its current state it would be laughed at. I imagine they need to improve quality massively before it’s viable to release.
But didn't Zuck say it will replace junior to mid level engineers on Joe Rogan podcast or something ?
Re: Meta Keeps Delaying the Release of Its New AI Model to Developers
#23I’ve used it at Meta. It’s very bad, if they released it in its current state it would be laughed at. I imagine they need to improve quality massively before it’s viable to release.
Re: Meta Keeps Delaying the Release of Its New AI Model to Developers
#24I’ve used it at Meta. It’s very bad, if they released it in its current state it would be laughed at. I imagine they need to improve quality massively before it’s viable to release.
This is what I suspected. Wang was a generationally bad hire. He has Meta SWEs making $250k+/year labeling data in AAI. He has exactly one move and it's this: https://i.imgflip.com/atotpp.jpg
I’m not sure the incentives are really aligned when you’re pouring that much cash and liquid RSUs at someone on normal vesting schedules. News stories of some of the acquisitions state that there are engineers in Meta’s AI organisation clearing 8 figures of compensation. If you didn’t think the strategy was successful, it’s rational (if not very principled) to continue to make excuses as to why until the gravy train stops and then use that to fund your retirement and the things you’d want to do instead.
Re: Meta Keeps Delaying the Release of Its New AI Model to Developers
#25Earlier quoted context omitted.
Seemed to work when it comes to selling ads. I'm thinking training LLMs is harder than anthropic and openai make it look
I'm guessing both openai and anthropic have transitioned to prompt magic and fine tuning rather than try to keep building LLMs at scale. The fact that QWEN and other models are impressive, small and perfectly suitable for most work means every dollar you're spending on trying to train larger models is a losing prop.
You probably don’t know how smaller models are trained then. Most of them are knowledge distilled or trained using data generated from larger models. If larger models are stopped there is no magical way smaller models will keep getting better.
Re: Meta Keeps Delaying the Release of Its New AI Model to Developers
#26Completely forgot that Meta was doing AI (and certainly spending billions doing so). They've got a lot of money, but are far behind on experience, talent, technology, and infrastructure.
Don't forget llama.cpp came about when meta released the weights to their LLaMa LLM. They've been in the game for awhile, just not anywhere near the top of the score board since.
Re: Meta Keeps Delaying the Release of Its New AI Model to Developers
#27Earlier quoted context omitted.
I'm guessing both openai and anthropic have transitioned to prompt magic and fine tuning rather than try to keep building LLMs at scale. The fact that QWEN and other models are impressive, small and perfectly suitable for most work means every dollar you're spending on trying to train larger models is a losing prop.
> every dollar you're spending on trying to train larger models is a losing prop You probably don’t know how smaller models are trained then. Most of them are knowledge distilled or trained using data generated from larger models. If larger models are stopped there is no magical way smaller models will keep getting better.
Re: Meta Keeps Delaying the Release of Its New AI Model to Developers
#28Earlier quoted context omitted.
> every dollar you're spending on trying to train larger models is a losing prop You probably don’t know how smaller models are trained then. Most of them are knowledge distilled or trained using data generated from larger models. If larger models are stopped there is no magical way smaller models will keep getting better.
you're arguing with capitalism not science or engineering.