Earlier quoted context omitted.
No, it's more like sharding of parameters. There's no understandable distinction between the experts.
I understand they're only optimizing for load distribution, but have people been trying to disentangle what the the various experts learn?
The Llama 4 herd
551–560 of 695 posts
Re: The Llama 4 herd
#552Earlier quoted context omitted.
No it is not. Right leaning opinions are heavily censored and shunned in all major publishing platforms that bots can scrape. For example, before Trump, if you contested the utterly normal common sense and scientifically sound idea that a trans woman is still a man, you would be banned - therefore, people with common sense will simply disengage, self-censor and get on with life.
Maybe because that position is both scientifically and morally unsound and if held strongly will lead to dehumanization and hate, attributes we should prevent any LLM from having.
It’s not immoral to recognize that you and your family and most of the people you know are split between penis and vagina.
It is immoral to police thoughts you disagree with. Believing race exists leads to dehumanization and hate. Maybe skin color doesn’t exist next? It’s just a representation with utility of similar feature/genetic groups that happened to evolve under similar environmental conditions. Is this scientifically unsound also?
Re: The Llama 4 herd
#553"It’s well-known that all leading LLMs have had issues with bias—specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet." Perhaps. Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. It's a simpler explanation.
Re: The Llama 4 herd
#554Earlier quoted context omitted.
Yeah truth itself is a bias. The idea of being unbiased doesn’t make sense.
I’ve seen more of this type of rhetoric online in the last few years and find it very insidious. It subtly erodes the value of objective truth and tries to paint it as only one of many interpretations or beliefs, which is nothing more than a false equivalence. The concept of being unbiased has been around for a long time, and we’re not going to throw it away just because a few people disagree with the premise.
Re: The Llama 4 herd
#555"It’s well-known that all leading LLMs have had issues with bias—specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet." Perhaps. Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. It's a simpler explanation.
Calling facts "playing into the leftists' agenda" is a problem of our shared political compass.
LLMs and humans need to do more work to implement doublethink, i.e. claiming non-truths and actually believing them to fit with a right-wing crowd for the sake of survival in it.
Re: The Llama 4 herd
#556Earlier quoted context omitted.
Call me crazy, but I don't want an AI that bases its reasoning on politics. I want one that is primarily scientific driven, and if I ask it political questions it should give me representative answers. E.g. "The majority view in [country] is [blah] with the minority view being [bleh]." I have no interest in "all sides are equal" answers because I don't believe all information is equally informative nor equally true.
The current crop of AIs can't do science though, they are disconnected from the physical world and can't test hypothesis or gather data.
Re: The Llama 4 herd
#557Model training observations from both Llama 3 and 4 papers: Meta’s Llama 3 was trained on ~16k H100s, achieving ~380–430 TFLOPS per GPU in BF16 precision, translating to a solid 38 - 43% hardware efficiency [Meta, Llama 3]. For Llama 4 training, Meta doubled the compute, using ~32K H100s and switched to FP8 precision. Despite the precision gain, observed efficiency dropped to about 19.7%, with GPUs delivering ~390 TF…
Even though it may not suitable for (existing) hardware impl, it may be advantageous in other place for example in learning rate speed.
Re: The Llama 4 herd
#558Earlier quoted context omitted.
This has been the case for a while now. 3090 hoarders were always just doing it for street cred or whatever, no way these guys are computing anything of actual value. Tenstorrent is on fire, though. For small businesses this is what matters. If 10M context is not a scam, I think we'll see SmartNIC adoption real soon. I would literally long AMD now because their Xilinx people are probably going to own the space real s…
> Infiniband is cool and all, but it's also stupid and their scale-out strategy is non-existent. god I love this website.
Re: The Llama 4 herd
#559Re: The Llama 4 herd
#560Earlier quoted context omitted.
This has been the case for a while now. 3090 hoarders were always just doing it for street cred or whatever, no way these guys are computing anything of actual value. Tenstorrent is on fire, though. For small businesses this is what matters. If 10M context is not a scam, I think we'll see SmartNIC adoption real soon. I would literally long AMD now because their Xilinx people are probably going to own the space real s…
Not a hoarder per-se but I bought a 24GB card on the secondary market. My privacy is valuable. I'm okay being a half-step or full-step behind in LLM or image diffusion if it means my data never leaves my machine.