How are Maverick and Scout distilled from Behemoth if the latter is not done training? Do they distill from some intermediate, "good enough" snapshot?
The Llama 4 herd
611–620 of 695 posts
Re: The Llama 4 herd
#612Earlier quoted context omitted.
Llama 4 Scout, Maximum context length: 10M tokens. This is a nice development.
Is the recall and reasoning equally good across the entirety of the 10M token window? Cause from what I've seen many of those window claims equate to more like a functional 1/10th or less context length.
Now maybe this is more a lack of instruction following than context length but the fact that it works at first and then starts going downhill quickly makes me wary about how much it will pay attention to other details further back in the context.
Re: The Llama 4 herd
#613Earlier quoted context omitted.
not everything an LLM is prompted for is comedy additionally, infantilizing entire groups of people is an ongoing criticism of the left by many groups of minorities, women, and the right. which is what you did by assuming it is “punching down”. the beneficiaries/subjects/victims of this infantilizing have said its not more productive than what overt racists/bigots do, and the left chooses to avoid any introspection o…
The leftist coddling crusades are just a different form of dominance over minorities. It absolutely is bigotry and sense of superiority driving it. That said, it would take one incredible therapist to get them to realize it.
I’ve never seen greater confusion in my life from otherwise well adjusted people.
“Self interest” is the go to term. “They’re [an amorphous group all in a single socioeconomic bracket] voting against their self interest”.
the form of dominance is very apparent but it seems like that crowd is completely blind to it, they're saying “here are the prepackaged things your kind can vote for, leave fiscal foreign and monetary policy to the white man. it is impossible for you to be in a position where those matters are relevant to you and may have you evaluating parties based on those factors. stick with the availability of elective surgeries like we said”
The left in the US manifests as the Democrat party, that party will be better off when they realize their constituents don’t really like them and are not that liberal. They're just more cautious of some people on the right.
Re: The Llama 4 herd
#614Earlier quoted context omitted.
Indeed, one of the notable things about LLMs is that the text they output is morally exemplary. This is because they are consistent in their rules. AI priests will likely be better than the real ones, consequently.
Quite the opposite. You can easily get a state of the art LLM to do a complete 180 on its entire moral framework with a few words injected in the prompt (and this very example demonstrates exactly that). It is very far from logically or ethically consistent. In fact it has no logic and ethics at all. Though if we did get an AI priest it would be great to absolve all your sins with some clever wordplay.
Re: The Llama 4 herd
#615Earlier quoted context omitted.
Ah, yes, the often used, peer-reviewed, expert-backed source of just listing random things. Thank you.
If you were looking for truth you wouldn’t reply like this. I’m not going to do an hour of work to carefully cite this for you, but it’s true nonetheless.
>If you were looking for truth
Except, with this, I don’t expect you to.
Re: The Llama 4 herd
#616Model training observations from both Llama 3 and 4 papers: Meta’s Llama 3 was trained on ~16k H100s, achieving ~380–430 TFLOPS per GPU in BF16 precision, translating to a solid 38 - 43% hardware efficiency [Meta, Llama 3]. For Llama 4 training, Meta doubled the compute, using ~32K H100s and switched to FP8 precision. Despite the precision gain, observed efficiency dropped to about 19.7%, with GPUs delivering ~390 TF…
Never trained a model, but the precision confused me as I've never considered how many bits should be reserved for exponent/mentisa. Has anyone architected a model(somehow) such that it has a free hand at using the give bits / choosing the type, or changed types from layer to layer, I mean surely when training for example vision models the first layers deal with the "big(yet simpler) picture"(light/dark, lines etc) w…
So between these four you honestly cover _most_ of the desired solution space: e.g. it's hard to imagine wanting to give up more of the mantissa than you already do on E5M2, while E4M3 is already at the lower bound of dynamic range before you need to start giving up IEEE compatability (which can definitely be a pain). There's some room left at the fp16 level but in practice bf16 was already designed for use in neural networks, so in practice people are happy using it for training and then leaving inference to fp16 (which has higher precision).
The only thing that's missing is support for more esoteric formats, e.g. fp4 (E2M1, E3M0) and maybe packed ternary.
Re: The Llama 4 herd
#617It's interesting that there are no reasoning models yet, 2.5 months after DeepSeek R1. It definitely looks like R1 surprised them. The released benchmarks look good. Large context windows will definitely be the trend in upcoming model releases. I'll soon be adding a new benchmark to test this more effectively than needle-in-a-haystack (there are already a couple of benchmarks that do that). All these models are very…
Re: The Llama 4 herd
#618It's interesting that there are no reasoning models yet, 2.5 months after DeepSeek R1. It definitely looks like R1 surprised them. The released benchmarks look good. Large context windows will definitely be the trend in upcoming model releases. I'll soon be adding a new benchmark to test this more effectively than needle-in-a-haystack (there are already a couple of benchmarks that do that). All these models are very…
But if the final result is of high enough quality, who cares about reasoning? It’s a trick to get the quality higher, at the cost of tokens and latency.
Re: The Llama 4 herd
#619"It’s well-known that all leading LLMs have had issues with bias—specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet." Perhaps. Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. It's a simpler explanation.
> Or, maybe, "leaning left" by the standards of Zuck et al. is more in alignment with the global population. The global population would be considered far-right by american standards. Particularly on LGBTQ matters and racism.
LGBTQ matters have varying degrees of acceptance around the world and Europe and the collective west are in front of it all, but that downplays the fact that LGBTQ acceptance has been rising nearly everywhere in the world with the exception of fundamentalist religious states.
Re: The Llama 4 herd
#620Earlier quoted context omitted.
"Honestly" and "literally" are now used in English for emphasis. I dislike this, but it's the current reality. I don't think there's any way to get back to only using them with their original meanings.
I don't think anyone needs to change their language. I understand that it's a common way to indicate candor, but it's hilariously inappropriate for a computer to say "some times I might lie to you to save your feelings, but this time, you really are ugly and you need to know."
Of course if you tink of the computer as a person you get strange results. A compiler error isn't the compiler telling me anything. It's the compiler writer telling me something. So a compiler error might contain a joke, and the joke might make sense, although obviously computers and compilers don't have a sense of humour.