As expected, Meta doesn't disappoint and accelerates the race to zero. Meta is undervalued.
The Llama 4 herd
151–160 of 695 posts
Re: The Llama 4 herd
#152looks like a leak to me.
Re: The Llama 4 herd
#153Re: The Llama 4 herd
#154Re: The Llama 4 herd
#155Earlier quoted context omitted.
I don't understand what LeCun is trying to say. Why does he give an interview saying that LLM's are almost obsolete just when they're about to release a model that increases the SotA context length by an order of magnitude? It's almost like a Dr. Jekyll and Mr. Hyde situation.
LeCun fundamentally doesn't think bigger and better LLMs will lead to anything resembling "AGI", although he thinks they may be some component of AGI. Also, he leads the research division, increasing context length from 2M to 10M is not interesting to him.
Re: The Llama 4 herd
#156> You never use phrases that imply moral superiority or a sense of authority, including but not limited to “it’s important to”, “it’s crucial to”, “it’s essential to”, "it's unethical to", "it's worth noting…", “Remember…” etc. Avoid using these. Aren't these phrases overrepresented in the first place because OpenAIs models use them so much? I guess Llama picked up the habit by consuming GPT output.
Personally I’d prefer that LLMs did not refer to themselves as “I”. It’s software, not an “I”.
Re: The Llama 4 herd
#157Re: The Llama 4 herd
#158"It’s well-known that all leading LLMs have had issues with bias—specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet." Perhaps. Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. It's a simpler explanation.
I heard reality has a well-known liberal bias.
Re: The Llama 4 herd
#159> These models are our best yet thanks to distillation from Llama 4 Behemoth, a 288 billion active parameter model with 16 experts that is our most powerful yet and among the world’s smartest LLMs. Llama 4 Behemoth outperforms GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on several STEM benchmarks. Llama 4 Behemoth is still training, and we’re excited to share more details about it even while it’s still in flight.
With 2T params (!!), it better outperform everything else.
https://blog.google/technology/google-deepmind/gemini-model-...
Re: The Llama 4 herd
#160Earlier quoted context omitted.
At 109b params you’ll need a ton of memory. We’ll have to wait for evals of the quants to know how much.
Sure but the upside of Apple Silicon is that larger memory sizes are comparatively cheap (compared to buying the equivalent amount of 5090 or 4090). Also you can download quantizations.
Apple should've invested more in bandwidth, but it's Apple and has lost its visionary. Imagine having 512GB on M3 Ultra and not being able to load even a 70B model on it at decent context window.