Live data from Hacker News

The Llama 4 herd

ai.meta.com

631–640 of 695 posts

Re: The Llama 4 herd

#631

Earlier quoted context omitted.

> Any position is a bias. A flat earther would consider a round-earther biased. That is not what a religion is. > Secular Americans are annoying because they believe they don't have one Why is that a problem to you? > and instead think they're just "good people", calling those who break their core values "bad people". No, not really. Someone is not good or bad because you agree with them. Even a religious person can…

> Why is that a problem to you? Because seculars/athiests often believe that they're superior to the "stupid, God-believing religious" people, since their beliefs are obviously based on "pure logic and reason". Yet, when you boil down anyone's value system to its fundamental essence, it turns out to always be a religious-like belief. No human value is based on pure logic, and it's annoying to see someone pretend othe…

People have value systems, yes. What's "boiling down" a value system?

You don't get to co-opt everybody as cryptically religious just because they have values.

Re: The Llama 4 herd

#632

Earlier quoted context omitted.

> Any position is a bias. A flat earther would consider a round-earther biased. That’s bollocks. The Earth is measurably not flat. You start from a position of moral relativism and then apply it to falsifiable propositions. It’s really not the same thing. Some ideas are provably false and saying that they are false is not "bias".

Dice are considered "biased" if not all sides have equal probability, even if that's literally true. When you look up the definition of bias you see "prejudice in favor of or against one thing, person, or group compared with another, usually in a way considered to be unfair." So the way we use the word has an implication of fairness to most people, and unfortunately reality isn't fair. Truth isn't fair. And that's wh…

Right. My point is that there are things we can argue about. "Is it better to have this road here or to keep the forest?", for example. Reasonable people can argue differently, and sensibility is important. Some would be biased towards business and economy, and others would be biased towards conservation. Having these debates in the media is helpful, even if you disagree.

But "is the Earth flat?" is no such question. Reasonable people cannot disagree, because the Earth is definitely not flat. Pretending like this is a discussion worth having is not being impartial, it’s doing a disservice to the audience.

Re: The Llama 4 herd

#633

Earlier quoted context omitted.

All right, let's say that the baseline is "what is true". Then bias is departure from the truth. That sounds great, right up until you try to do something with it. You want your LLM to be unbiased? So you're only going to train it on the truth? Where are you going to find that truth? Oh, humans are going to determine it? Well, first, where are you going to find unbiased humans? And, second, they're going to curate al…

The definition of the word has no responsibility to your opinion of it as an epistemology. Also, you're just complaining about the difficulty of determining what is true. That's a separate problem, isn't it?

If we had an authoritative way of determining truth, then we wouldn't have the problem of curating material to train an LLM on. So no, I don't think it's a separate problem.

Re: The Llama 4 herd

#634

Earlier quoted context omitted.

The definition of the word has no responsibility to your opinion of it as an epistemology. Also, you're just complaining about the difficulty of determining what is true. That's a separate problem, isn't it?

If we had an authoritative way of determining truth, then we wouldn't have the problem of curating material to train an LLM on. So no, I don't think it's a separate problem.

Again, the word "bias" and its definition exists outside the comparatively narrow concern of training LLMs.

Re: The Llama 4 herd

#636

Llama 4 Maverick scored 16% on the aider polyglot coding benchmark [0]. 73% Gemini 2.5 Pro (SOTA) 60% Sonnet 3.7 (no thinking) 55% DeepSeek V3 0324 22% Qwen Max 16% Qwen2.5-Coder-32B-Instruct 16% Llama 4 Maverick [0] https://aider.chat/docs/leaderboards/?highlight=Maverick

Did they not target code tasks for this LLM, or is it genuinely that bad? Pretty embarrassing when your shiny new 400B model barely ties a 32B model designed to be run locally. Or maybe is this a strong indication that smaller, specialized LLMs have much more potential for specific tasks than larger, general purpose LLMs.

Re: The Llama 4 herd

#637
post #450

Earlier quoted context omitted.

BTW, I'd love to see a large model designed from scratch for efficient local inference on low-memory devices. While current MoE implementations are tuned for load-balancing over large pools of GPUs, there is nothing stopping you tuning them to only switch expert once or twice per token, and ideally keep the same weights across multiple tokens. Well, nothing stopping you, but there is the question of if it will actual…

Intuitively it feels like there ought to be significant similarities between expert layers because there are fundamentals about processing the stream of tokens that must be shared just from the geometry of the problem. If that's true, then identifying a common abstract base "expert" then specialising the individuals as low-rank adaptations on top of that base would mean you could save a lot of VRAM and expert-swappin…

Yes, Deepseek introduced this optimisation of a common base "expert" that's always loaded. Llama 4 uses it too.

Re: The Llama 4 herd

#638

Earlier quoted context omitted.

If we had an authoritative way of determining truth, then we wouldn't have the problem of curating material to train an LLM on. So no, I don't think it's a separate problem.

Again, the word "bias" and its definition exists outside the comparatively narrow concern of training LLMs.

So? The smaller problem is solved by solving the larger problem. So, not separate problems.

You seem to have a larger point or position or something that you're hinting at. Would you stop being vague, and actually state what's on your mind?

Re: The Llama 4 herd

#639

Earlier quoted context omitted.

It is not and has never been half. 2024 voter turnout was 64%

> It is not and has never been half. 2024 voter turnout was 64% He said half of voters, those who didn't vote aren't voters.

When I replied, the comment said "Americans", per the edit

Re: The Llama 4 herd

#640
post #560

Earlier quoted context omitted.

If you really were serious about privacy, you wouldn't put yourself at disadvantage with a locked-down six-year out of date card. Tenstorrent Blackhole exists now, btw.

I think it's disingenuous to suggest they're putting themselves at a disadvantage with an RTX 3090, especially in a comparison to an inferior product that isn't even shipping yet. RTX 3090: 24GB RAM, 936.2GB/s bandwidth Tenstorrent p150a: 32GB RAM, 512GB/s bandwidth an extra 8GB of ram isn't worth nearly halving memory bandwidth.

> inferior product

Tenstorrent p300 is coming at 64 GB and 1 Tbps but that's not the point; even p150a with plenty of bandwidth (512 GB/s is fine for inference) and four 800G ports. But hardware is not the problem: even if they had the hardware, they wouldn't know what to do with it. Privacy is a hobby to most people, making you feel good.

Post reply on HN