Live data from Hacker News

The Llama 4 herd

ai.meta.com

651–660 of 695 posts

Re: The Llama 4 herd

#651
post #604

Earlier quoted context omitted.

I think it's disingenuous to suggest they're putting themselves at a disadvantage with an RTX 3090, especially in a comparison to an inferior product that isn't even shipping yet. RTX 3090: 24GB RAM, 936.2GB/s bandwidth Tenstorrent p150a: 32GB RAM, 512GB/s bandwidth an extra 8GB of ram isn't worth nearly halving memory bandwidth.

Or how about https://www.notebookcheck.net/Way-to-run-DeepSeek-s-671B-AI-... 768 GiB for $6000.

I might be filing for bankruptcy soon so I'm definitely stuck with what I've got.

Re: The Llama 4 herd

#652

General overview below, as the pages don't seem to be working well Llama 4 Models: - Both Llama 4 Scout and Llama 4 Maverick use a Mixture-of-Experts (MoE) design with 17B active parameters each. - They are natively multimodal: text + image input, text-only output. - Key achievements include industry-leading context lengths, strong coding/reasoning performance, and improved multilingual capabilities. - Knowledge cuto…

If their knowledge cutoff is 8 months ago, then how on earth does Grok know things that happened yesterday?

I would really love to know that.

Re: The Llama 4 herd

#653

I guess I have to say thank you Meta? A somewhat sad rant below. Deepseek starts a toxic trend of providing super, super large MoE. And MoE is famous for being parameter-inefficient, which is unfriendly to normal consumer hardware with limited vram. The super large size of LLM also disables nearly every people from doing meaningful development on these models. R1-1776 is the only fine-tune variation of R1 that makes…

Have you heard of the bitter lesson? Bigger means better in Neural Networks.

Yeah. I know the bitter lesson.

For neutral networks, on one hand, larger size generally indicates higher performance upper limit. On the other hand, you really have to find ways to materialize these advantages over small models, or larger size becomes a burden.

However, I'm talking about local usage of LLMs instead of production usage, which is severely limited by GPUs with low VRAM. You literally cannot run LLMs beyond a specific size.

Re: The Llama 4 herd

#655

Earlier quoted context omitted.

At 4 bit quant (requires 64GB) the price of Mac (4.2K) is almost exactly the same as 2x5090 (provided we will see them in stock). But 2x5090 have 6x memory bandwidth and probably close to 50x matmul compute at int4.

2.8k-3.6k for a 64gb-128gb mac studio (m3 max).

If you go a gen or two back, you can get 3x3090 for the same price.

Re: The Llama 4 herd

#656

Earlier quoted context omitted.

Maybe because that position is both scientifically and morally unsound and if held strongly will lead to dehumanization and hate, attributes we should prevent any LLM from having.

You’re very confident in your opinions. It’s not immoral to recognize that you and your family and most of the people you know are split between penis and vagina. It is immoral to police thoughts you disagree with. Believing race exists leads to dehumanization and hate. Maybe skin color doesn’t exist next? It’s just a representation with utility of similar feature/genetic groups that happened to evolve under similar…

Not everyone has either or, some even have both

Re: The Llama 4 herd

#657
post #68

"It’s well-known that all leading LLMs have had issues with bias—specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet." Perhaps. Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. It's a simpler explanation.

It's a bit of both, but the point holds. Pre-Musk Twitter and Reddit are large datasources and they leaned hard-left, mostly because of censorship.

Re: The Llama 4 herd

#658
post #68

"It’s well-known that all leading LLMs have had issues with bias—specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet." Perhaps. Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. It's a simpler explanation.

> is more in alignment with the global population

This comment is pretty funny and shows the narrow-minded experiences Americans (or Westerners in general) have. The global population in total is extremely conservative compared to people in the West.

Re: The Llama 4 herd

#659
post #230

Earlier quoted context omitted.

Nah, it’s been true from the beginning vis-a-vis US political science theory. That is, if you deliver something like https://www.pewresearch.org/politics/quiz/political-typology... To models from GPT-3 on you get highly “liberal” per Pew’s designations. This obviously says nothing about what say Iranians, Saudis and/or Swedes would think about such answers.

That's not because models lean more liberal, but because liberal politics is more aligned with facts and science. Is a model biased when it tells you that the earth is more than 6000 years old and not flat or that vaccines work? Not everything needs a "neutral" answer.

> but because liberal politics is more aligned with facts and science

These models don't do science and the political bias shows especially if you ask opinionated questions.

Re: The Llama 4 herd

#660

Earlier quoted context omitted.

So google Gemini was creating black Vikings because of facts?

Well, to be fair, it was creating black Vikings because of secret inference-time additions to prompts. I for one welcome Vikings of all colors if they are not bent on pillage or havoc

> secret inference-time additions to prompts

Which were politically biased, in turn making the above assumption true.

Post reply on HN