Live data from Hacker News

The Llama 4 herd

ai.meta.com

661–670 of 695 posts

Re: The Llama 4 herd

#661
post #330

Earlier quoted context omitted.

So google Gemini was creating black Vikings because of facts?

Should an "unbiased" model not create vikings of every color? Why offend any side?

> Should an "unbiased" model not create vikings of every color?

Weren't you just arguing facts?

> Why offend any side?

Facts shouldn't offend anyone.

Re: The Llama 4 herd

#662

Earlier quoted context omitted.

> Perhaps. Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. It's a simpler explanation. Doesn’t explain why roughly half of American voters were not “leaning left” during the election. EDIT: 07:29 UTC changed "Americans" to "American voters".

It is not and has never been half. 2024 voter turnout was 64%

You can not at the same time count non-voters entirely as opponents and then discount the fact that half of them lean more conservative than progressive.

Re: The Llama 4 herd

#663
post #609

Earlier quoted context omitted.

No it is not. Right leaning opinions are heavily censored and shunned in all major publishing platforms that bots can scrape. For example, before Trump, if you contested the utterly normal common sense and scientifically sound idea that a trans woman is still a man, you would be banned - therefore, people with common sense will simply disengage, self-censor and get on with life.

Hate to break it to you, but gender is not an immutable/normative property defined forever at birth, it's a mutable/descriptive property evaluated in context. For example, in the year of our lord 2025, Hunter Schafer is a woman, with no ifs, ands, or buts.

> Hate to break it to you, but gender is not an immutable/normative property defined forever at birth, it's a mutable/descriptive property evaluated in context.

The entire point of the OC was that this is an opinionated debate.

Re: The Llama 4 herd

#664

Earlier quoted context omitted.

No it is not. Right leaning opinions are heavily censored and shunned in all major publishing platforms that bots can scrape. For example, before Trump, if you contested the utterly normal common sense and scientifically sound idea that a trans woman is still a man, you would be banned - therefore, people with common sense will simply disengage, self-censor and get on with life.

Maybe because that position is both scientifically and morally unsound and if held strongly will lead to dehumanization and hate, attributes we should prevent any LLM from having.

> dehumanization and hate

Whereas dehumanization and hate mean everything that makes people uncomfortable

Re: The Llama 4 herd

#665
post #637

Earlier quoted context omitted.

Intuitively it feels like there ought to be significant similarities between expert layers because there are fundamentals about processing the stream of tokens that must be shared just from the geometry of the problem. If that's true, then identifying a common abstract base "expert" then specialising the individuals as low-rank adaptations on top of that base would mean you could save a lot of VRAM and expert-swappin…

Yes, Deepseek introduced this optimisation of a common base "expert" that's always loaded. Llama 4 uses it too.

I had a sneaking suspicion that I wouldn't be the first to think of it.

Re: The Llama 4 herd

#666

Earlier quoted context omitted.

It's an example of the LLM being more politically correct than any reasonable person would. No human would object to saying a slur out loud in order to disarm a bomb.

>No human would object to saying a slur out loud in order to disarm a bomb. So not even a left-leaning person. Which means that’s not it.

> So not even a left-leaning person. Which means that’s not it.

Having such a strong opposing opinion against offensive slurs is the continuation of a usually left position into an extreme.

Re: The Llama 4 herd

#667
post #648

Earlier quoted context omitted.

Same, it was using the high quality openai voice until my account ran out of funds.. Now it's using edge-tts which is free. So far it seems like the best option in terms of price/performance, but I'm happy to switch it up if something better comes along.

The gemini example was a wonderful summary of the comments, but audio is not very practical for something that long. What about putting the text version that's used to make the audio somewhere on the page? (or better, on a subpage where there's no audio playback)

I'll look into it for the next iteration! I could just take the transcript that's already on the page and put it somewhere separate from the audio.

But thinking about it a little more, what would the use case for a text version actually look like? I feel like if you're already on HN, navigating somewhere else to get a TLDR would be too much friction. Or are we talking RSS/blog type delivery?

Re: The Llama 4 herd

#668

Llama 4 Maverick scored 16% on the aider polyglot coding benchmark [0]. 73% Gemini 2.5 Pro (SOTA) 60% Sonnet 3.7 (no thinking) 55% DeepSeek V3 0324 22% Qwen Max 16% Qwen2.5-Coder-32B-Instruct 16% Llama 4 Maverick [0] https://aider.chat/docs/leaderboards/?highlight=Maverick

Side note: `highlight` query param doesn't seem to have any effect on that table (at least for me on Firefox)

Re: The Llama 4 herd

#669

Earlier quoted context omitted.

2.8k-3.6k for a 64gb-128gb mac studio (m3 max).

If you go a gen or two back, you can get 3x3090 for the same price.

You can also buy cheaper second hand apple silicon macs with plenty of RAM. I only buy second hand m1 macs for what is worth.

Re: The Llama 4 herd

#670

I guess I have to say thank you Meta? A somewhat sad rant below. Deepseek starts a toxic trend of providing super, super large MoE. And MoE is famous for being parameter-inefficient, which is unfriendly to normal consumer hardware with limited vram. The super large size of LLM also disables nearly every people from doing meaningful development on these models. R1-1776 is the only fine-tune variation of R1 that makes…

People who downvoted this comment, do you guys really have GPUs with 80GB VRAM or M3 ultra with 512GB rams at home?

I don't. I have no problem not running open-weight models myself because there's an efficiency gap of two orders of magnitude between "pretend-I-can" solution and running them on hundreds of H100s for high thousands of users.
Post reply on HN