Can someone explain to me please why Meta doesn't create subject specific versions of their LLMs such as one that knows only about computer programming, computers, hardware software. I would have imagined such a thing would be smaller and thus run on smaller configurations. But since I am only a layman maybe someone can tell me why this isn't the case?
Everything we announced at our first LlamaCon
41–50 of 122 posts
Re: Everything we announced at our first LlamaCon
#42Earlier quoted context omitted.
>There is a potential world where Meta [is]... literally building smart homes... while preserving privacy and building trust I'm earnestly not sure what Meta are less qualified for. Building physical homes or building privacy & trust.
Visit SE Asia sometime and you'll experience a very different sentiment. Hundreds of millions of people rely on Meta to provide valuable services every day, some of them borderline essential. This is undebatable. The outsized public hatred toward Meta is almost entirely driven by a bureaucratic, anti-technology Europe (that has finally realized that their overstepping is hurting their future) and a US political insti…
Re: Everything we announced at our first LlamaCon
#43Earlier quoted context omitted.
>There is a potential world where Meta [is]... literally building smart homes... while preserving privacy and building trust I'm earnestly not sure what Meta are less qualified for. Building physical homes or building privacy & trust.
Visit SE Asia sometime and you'll experience a very different sentiment. Hundreds of millions of people rely on Meta to provide valuable services every day, some of them borderline essential. This is undebatable. The outsized public hatred toward Meta is almost entirely driven by a bureaucratic, anti-technology Europe (that has finally realized that their overstepping is hurting their future) and a US political insti…
That is the economic structure of their business model.
Now juice that model with $ billions of revenue and $ trillions in potential market cap for shareholders, who demand double digit percentage growth per year.
That defines the scale of available resources to drive the business model forward.
This is a machine designed to scale up and maximally leverage seemingly small conflicts of interest into a global monster that feeds on mental and social decay.
——
Of course, it benefits Facebook and customers to mix in as much genuine side products and services with real value as possible.
But that only wedges the destructive core into individual lives and society even more.
Now add AI algorithms to their core competencies of surveillance integration and psychological manipulation, and to the side value honey features.
We are getting Stockholm’ed and stewed in a lot of high walled slow cookers these days.
Re: Everything we announced at our first LlamaCon
#44Lmao why are they doing LlamaCon, a convention with a subpar product?
The problem, in my opinion, is that MZ/CC/AA-D, are feeling that they have to be releasing models of some flavor every month to stay competitive.
And when you have the rest of the company planning to throw you a on-stage party to announce whatever next model, and the venue and guests are paid for, you're gonna have the show whether the content is good or not.
Llama program right now is "we must go faster." But without a clear product direction or niche that they're trying to build towards. Very little is said no to. Just be the best at everything. And they started from behind, how can you think you're gonna catch up to 1-2 year head start, just with more people? The line they want to believe is "the best LLM, not just the best OSS LLM".
Because of the constant pressure to release something every month (nearly, but not a huge exaggeration), and the product direction coming from MZ himself, the team is not really great at anything. There is a huge apparatus of people working on it, yet half of it or more, I believe, is baggage required because of what Meta is.
I guess we'll see how long this can be maintained.
Re: Everything we announced at our first LlamaCon
#450. Introducing Llama API in preview
This one is good but not centre stage worthy. Other [closed] models have been offering this for a long time.
1. Fast inference with Llama API
How fast? and how must faster than others? This section talks about latency and there's absolutely no numbers in this section!
2. New Llama Stack integrations
Speculations with 0 new integration. Llama Stack with NVIDIA had already been announced and then this section ends with '...others on new integrations that will be announced soon. Alongside our partners, we envision Llama Stack as the industry standard for enterprises looking to seamlessly deploy production-grade turnkey AI solutions.'
3. New Llama Protections and security for the open source community
This one is not only the best on this page, but is actually good with announcement of - Llama Guard 4, LlamaFirewall, and Llama Prompt Guard 2
4. Meet the Llama Impact Grant recipients
Sorry but neither the gross amount $1.5 million USD, nor the average $150K/recipients is anything significant at Facebook scale.
Re: Everything we announced at our first LlamaCon
#46Re: Everything we announced at our first LlamaCon
#47Feels like Meta is going to Cloud services business but in AI domain. They resisted entering cloud business for so long, with the success of AWS/Azure/GCP I think they are realizing they can't keep at the top only with social networks without owning a platform (hardware, cloud)
Re: Everything we announced at our first LlamaCon
#48Its all Ai all the time now though, not seen any mention of our reimagined future of floating heads hanging out together in quite some time.
Re: Everything we announced at our first LlamaCon
#49Earlier quoted context omitted.
What's the minimum GPU/NPU hardware and memory to run Qwen3 locally?
`model.safetensors` for Qwen3-0.6B is a single 1.5GB file. Qwen3-235B-A22B has 118 `.safetensors` files at 4GB each. There are a bunch of models and quants between those.
Re: Everything we announced at our first LlamaCon
#50Meta needs to stop open-washing their product. It simply is not open-source. The license for their precompiled binary blob (ie model) should not be considered open-source, and the source code (ie training process / data) isn’t available.