Live data from Hacker News

The Llama 4 herd

ai.meta.com

521–530 of 695 posts

Re: The Llama 4 herd

#521

I don't really understand how Scout and Maverick are distillations of Behemoth if Behemoth is still training. Maybe I missed or misunderstood this in the post? Did they distill the in-progress Behemoth and the result was good enough for models of those sizes for them to consider releasing it? Or is Behemoth just going through post-training that takes longer than post-training the distilled versions? Sorry if this is…

My understanding is that they have a base model checkpoint for Behemoth from pre-training.

This base model is not instruction-tuned so you can't use it like a normal instruction-tuned model for chatbots.

However, the base model can be distilled, and then the distilled model is post-trained to be instruction tuned, which can be released as a model for chatbots.

Re: The Llama 4 herd

#522

I guess I have to say thank you Meta? A somewhat sad rant below. Deepseek starts a toxic trend of providing super, super large MoE. And MoE is famous for being parameter-inefficient, which is unfriendly to normal consumer hardware with limited vram. The super large size of LLM also disables nearly every people from doing meaningful development on these models. R1-1776 is the only fine-tune variation of R1 that makes…

More on the accessibility problem, even a request from a Meta engineer was rejected. Is that normal?

See https://huggingface.co/spaces/meta-llama/README/discussions/...

Re: The Llama 4 herd

#523

I don't really understand how Scout and Maverick are distillations of Behemoth if Behemoth is still training. Maybe I missed or misunderstood this in the post? Did they distill the in-progress Behemoth and the result was good enough for models of those sizes for them to consider releasing it? Or is Behemoth just going through post-training that takes longer than post-training the distilled versions? Sorry if this is…

> Or is Behemoth just going through post-training that takes longer than post-training the distilled versions?

This is the likely main explanation. RL fine-tuning repeatedly switches between inference to generate and score responses, and training on those responses. In inference mode they can parallelize across responses, but each response is still generated one token at a time. Likely 5+ minutes per iteration if they're aiming for 10k+ CoTs like other reasoning models.

There's also likely an element of strategy involved. We've already seen OpenAI hold back releases to time them to undermine competitors' releases (see o3-mini's release date & pricing vs R1's). Meta probably wants to keep that option open.

Re: The Llama 4 herd

#524

Earlier quoted context omitted.

I follow a secular humanist moral system as best I can. I have tolerance for those who have tolerance for me. I grew up amongst fundamentalist christians and fundamentalist anything (christian, muslim, buddhist, whatever) leave a bad taste in my mouth. I don't care about your religion just don't try to force it on me or try to make me live by its moral system and you won't hear a peep out of me about what you're doin…

That's a fine attitude, but now you're describing your own beliefs rather than "the right" or "the left". Statistically, white people make more money than black people and men make more money than women and there are differences in their proportions in various occupations. This could be caused by cultural differences that correlate with race, or hormonal differences that cause behavioral differences and correlate wit…

While I believe there might be different explanations for the outcomes we observe I also believe that default hypothesis should be that there is racism and sexism. And there are facts (women were permitted to vote in the US like 100 years ago, and entered general workforce when?), observations (I saw sexism and racism at work) and general studies (I.e people have tendency to have biases among other things) to support that attributing differences to biology or whatever should be under very high scrutiny.

Re: The Llama 4 herd

#525
post #68

"It’s well-known that all leading LLMs have had issues with bias—specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet." Perhaps. Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. It's a simpler explanation.

> global population

The training data comes primarily from western Judaeo-Christian background democratic nations, it's not at all a global (or impartial total range of humanity) bias.

Re: The Llama 4 herd

#526
post #230

Earlier quoted context omitted.

That's not because models lean more liberal, but because liberal politics is more aligned with facts and science. Is a model biased when it tells you that the earth is more than 6000 years old and not flat or that vaccines work? Not everything needs a "neutral" answer.

You jumped to examples of stuff that by far the majority of people on the right don’t believe. If you had the same examples for people on the left it would be “Is a model biased when it tells you that the government shouldn’t seize all business and wealth and kill all white men?” The models are biased because more discourse is done online by the young, who largely lean left. Voting systems in places like Reddit make…

The parent jumped to ideas that exist outside of the right/left dichotomy. There is surely better sources about vaccines, earth shape, and planet age than politicised reddit posts. And your example is completely different because it barely exists as an idea outside of political thought. Its a tiny part of human thought.

Re: The Llama 4 herd

#527
post #63

Earlier quoted context omitted.

Llama 4 Scout, Maximum context length: 10M tokens. This is a nice development.

I don't think RAG will survive this time

This is only for the small model. The medium model is still at 1M (like Gemini 2.5)

Even if we could get the mid models to 10M, that's still a medium-sized repo at best. Repos size growth will also accelerate as LLMs generate more code. There's no way to catch up.

Re: The Llama 4 herd

#528
post #230

Earlier quoted context omitted.

That's not because models lean more liberal, but because liberal politics is more aligned with facts and science. Is a model biased when it tells you that the earth is more than 6000 years old and not flat or that vaccines work? Not everything needs a "neutral" answer.

> Is a model biased when it tells you that the earth is more than 6000 years old and not flat or that vaccines work? Not everything needs a "neutral" answer. That's the motte and bailey. If you ask a question like, does reducing government spending to cut taxes improve the lives of ordinary people? That isn't a science question about CO2 levels or established biology. It depends on what the taxes are imposed on, the…

> But in politics it does, which is that the right says yes and the left says no.

That’s not accurate, tax deductions for the poor is an obvious example. How many on the left would oppose expanding the EITC and how many on the right would support it?

Re: The Llama 4 herd

#529

Earlier quoted context omitted.

I find it impossible to discuss bias without a shared understanding of what it actually means to be unbiased - or at least, a shared understanding of what the process of reaching an unbiased position looks like. 40% of Americans believe that God created the earth in the last 10,000 years. If I ask an LLM how old the Earth is, and it replies ~4.5 billion years old, is it biased?

I've wondered if political biases are more about consistency than a right or left leaning. For instance, if I train a LLM only on right-wing sources before 2024, and then that LLM says that a President weakening the US Dollar is bad, is the LLM showing a left-wing bias? How did my LLM trained on only right-wing sources end up having a left-wing bias? If one party is more consistent than another, then the underlying l…

> because things like this are not naturally equal.

Really? Seems to me like no one has the singular line on reality, and everyone's perceptions are uniquely and contextually their own.

Wrong is relative: https://hermiene.net/essays-trans/relativity_of_wrong.html

But it seems certain that we're all wrong about something. The brain does not contain enough bits to accurately represent reality.

Re: The Llama 4 herd

#530

Earlier quoted context omitted.

Sure but the upside of Apple Silicon is that larger memory sizes are comparatively cheap (compared to buying the equivalent amount of 5090 or 4090). Also you can download quantizations.

At 4 bit quant (requires 64GB) the price of Mac (4.2K) is almost exactly the same as 2x5090 (provided we will see them in stock). But 2x5090 have 6x memory bandwidth and probably close to 50x matmul compute at int4.

2.8k-3.6k for a 64gb-128gb mac studio (m3 max).
Post reply on HN