Earlier quoted context omitted.
I think it's disingenuous to suggest they're putting themselves at a disadvantage with an RTX 3090, especially in a comparison to an inferior product that isn't even shipping yet. RTX 3090: 24GB RAM, 936.2GB/s bandwidth Tenstorrent p150a: 32GB RAM, 512GB/s bandwidth an extra 8GB of ram isn't worth nearly halving memory bandwidth.
Or how about https://www.notebookcheck.net/Way-to-run-DeepSeek-s-671B-AI-... 768 GiB for $6000.
The Llama 4 herd
651–660 of 695 posts
Re: The Llama 4 herd
#652General overview below, as the pages don't seem to be working well Llama 4 Models: - Both Llama 4 Scout and Llama 4 Maverick use a Mixture-of-Experts (MoE) design with 17B active parameters each. - They are natively multimodal: text + image input, text-only output. - Key achievements include industry-leading context lengths, strong coding/reasoning performance, and improved multilingual capabilities. - Knowledge cuto…
I would really love to know that.
Re: The Llama 4 herd
#653I guess I have to say thank you Meta? A somewhat sad rant below. Deepseek starts a toxic trend of providing super, super large MoE. And MoE is famous for being parameter-inefficient, which is unfriendly to normal consumer hardware with limited vram. The super large size of LLM also disables nearly every people from doing meaningful development on these models. R1-1776 is the only fine-tune variation of R1 that makes…
Have you heard of the bitter lesson? Bigger means better in Neural Networks.
For neutral networks, on one hand, larger size generally indicates higher performance upper limit. On the other hand, you really have to find ways to materialize these advantages over small models, or larger size becomes a burden.
However, I'm talking about local usage of LLMs instead of production usage, which is severely limited by GPUs with low VRAM. You literally cannot run LLMs beyond a specific size.
Re: The Llama 4 herd
#654Re: The Llama 4 herd
#655Earlier quoted context omitted.
At 4 bit quant (requires 64GB) the price of Mac (4.2K) is almost exactly the same as 2x5090 (provided we will see them in stock). But 2x5090 have 6x memory bandwidth and probably close to 50x matmul compute at int4.
2.8k-3.6k for a 64gb-128gb mac studio (m3 max).
Re: The Llama 4 herd
#656Earlier quoted context omitted.
Maybe because that position is both scientifically and morally unsound and if held strongly will lead to dehumanization and hate, attributes we should prevent any LLM from having.
You’re very confident in your opinions. It’s not immoral to recognize that you and your family and most of the people you know are split between penis and vagina. It is immoral to police thoughts you disagree with. Believing race exists leads to dehumanization and hate. Maybe skin color doesn’t exist next? It’s just a representation with utility of similar feature/genetic groups that happened to evolve under similar…
Re: The Llama 4 herd
#657"It’s well-known that all leading LLMs have had issues with bias—specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet." Perhaps. Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. It's a simpler explanation.
Re: The Llama 4 herd
#658"It’s well-known that all leading LLMs have had issues with bias—specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet." Perhaps. Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. It's a simpler explanation.
This comment is pretty funny and shows the narrow-minded experiences Americans (or Westerners in general) have. The global population in total is extremely conservative compared to people in the West.
Re: The Llama 4 herd
#659Earlier quoted context omitted.
Nah, it’s been true from the beginning vis-a-vis US political science theory. That is, if you deliver something like https://www.pewresearch.org/politics/quiz/political-typology... To models from GPT-3 on you get highly “liberal” per Pew’s designations. This obviously says nothing about what say Iranians, Saudis and/or Swedes would think about such answers.
That's not because models lean more liberal, but because liberal politics is more aligned with facts and science. Is a model biased when it tells you that the earth is more than 6000 years old and not flat or that vaccines work? Not everything needs a "neutral" answer.
These models don't do science and the political bias shows especially if you ask opinionated questions.
Re: The Llama 4 herd
#660Earlier quoted context omitted.
So google Gemini was creating black Vikings because of facts?
Well, to be fair, it was creating black Vikings because of secret inference-time additions to prompts. I for one welcome Vikings of all colors if they are not bent on pillage or havoc
Which were politically biased, in turn making the above assumption true.