Earlier quoted context omitted.
This was an idea that sounded somewhat silly until it was shown it worked. The idea is that you encourage through training a bunch of “experts” to diversify and “get good” at different things. These experts are say 1/10 to 1/100 of your model size if it were a dense model. So you pack them all up into one model, and you add a layer or a few layers that have the job of picking which small expert model is best for your…
If I have 5000 documents about A, and 5000 documents about B, do we know whether it's better to train one large model on all 10,000 documents, or to train 2 different specialist models and then combine them as you describe?
The Llama 4 herd
621–630 of 695 posts
Re: The Llama 4 herd
#622Re: The Llama 4 herd
#623Earlier quoted context omitted.
For a super ignorant person: Both Llama 4 Scout and Llama 4 Maverick use a Mixture-of-Experts (MoE) design with 17B active parameters each Those experts are LLM trained on specific tasks or what?
This was an idea that sounded somewhat silly until it was shown it worked. The idea is that you encourage through training a bunch of “experts” to diversify and “get good” at different things. These experts are say 1/10 to 1/100 of your model size if it were a dense model. So you pack them all up into one model, and you add a layer or a few layers that have the job of picking which small expert model is best for your…
Makes sense to compare apples with apples. Same compute amount, right? Or you are giving less time to MoE model and then feel like it underperforms. Shouldn't be surprising...
> These experts are say 1/10 to 1/100 of your model size if it were a dense model
Just to be correct, each layer (attention + fully connected) has it's own router and experts. There are usually 30++ layers. It can't be 1/10 per expert as there are literally hundreds of them.
Re: The Llama 4 herd
#624Earlier quoted context omitted.
Llama 4 Scout, Maximum context length: 10M tokens. This is a nice development.
Is the recall and reasoning equally good across the entirety of the 10M token window? Cause from what I've seen many of those window claims equate to more like a functional 1/10th or less context length.
Re: The Llama 4 herd
#625Can we somehow load these inside node.js? What is the easiest way to load them remotely? Huggingface Spaces? Google AI Studio? I am teaching a course on AI to non-technical students, and I wanted the students to have a minimal setup: which in this case would be: 1) Browser with JS (simple folder of HTML, CSS) and Tensorflow.js that can run models like Blazeface for face recognition, eye tracking etc. (available since…
You're not qualified to teach a course on AI if you're asking questions like that. Please don't scam students, they're naive and don't know better and you're predating.
Oh trust me, I am very upfront about what I know and do not know. My main background is in developing full stack web sites, apps, and using APIs. I have been using AI models since 2019, using Tensorflow.js in the browser and using APIs for years. I am not in the Python ecosystem, though, I don’t go deep into ML and don’t pretend to. I don’t spend my days with PyTorch, CUDA or fine-tuning models or running my own infrastructure.
Your comment sounds like “you don’t know cryptographyc if you have to ask basic questions about quantum-resistant SPHICS+ or bilinear pairings, do not teach a class on how to succeed in business using blockchain and crypto, you’re scamming people.”
Or in 2014: “if you don’t know how QUIC and HTTP/2 works and Web Push and WebRTC signaling, and the latest Angular/React/Vue/Svelte/… you aren’t qualified to teach business school students how to make money with web technology”.
It’s the classic engineering geek argument. But most people can make money without knowing the ins and outs of every single technology, every single framework. It is much more valuable to see what works and how to use it. Especially when the space changes week to week as I teach it. The stuff I teach in the beginning of the course (eg RAG) may be obsolete by the time the latest 10-million token model drops.
I did found an AI startup a few years ago and was one of the first to use OpenAI’s completions API to build bots for forums etc. I also work to connect deep tech to AI, to augment it: https://engageusers.ai/ecosystem.pdf
And besides — every time I start getting deep into how the models work, including RoPe and self—attention and transformer architecture, their eyes glaze over. They barely know the difference between a linear function wnd an activation function. At best I am giving these non-technical business students three things:
1) an intuition about how the models are trained, do inference and how jobs are submitted, to take the magic out of it. I showed them everything from LLMs to Diffusion models and GANs, but I keep emphasizing that the techniques are improving
2) how to USE the latest tools like bolt.new or lovable or opusclip etc.
3) do hands-on group projects to simulate working on a team and building a stack, that’s how I grade them. And for this I wanted to MINIMIZE what they need to install. LLaMa 4 for one GPU is the ticket!
Yeah so I was hoping the JS support was more robust, and asking HN if they knew of any ports (at least to WASM). But no, it’s firmly locked into PyTorch and CUDA for now. So I’m just gonna stick with Tensorflow for educational purposes, like people used Pascal or Ruby when teaching. I want to let them actually install ONE thing (node.js) and be able to run inferenfe in their browser. I want them to be able to USE the tools and build websites and businesses end-to-end, launch a business and have agents work for them.
Some of the places they engage the most is when I talk about AI and society, sustainability or regulations. That’s the cohort
But you can keep geeking out on low-level primitives. I remember writing my own 3D-persoective-correct-texturemapping engine and then GPUs came out. Carmack and others kept at it for a while, others moved on. You could make a lot of money in 3D games without knowing how texturemapping and lighting worked, and same goes for this.
PS: No thanks to you but I found what I was looking for myself in a few minutes. https://youtu.be/6LHNbeDADA4?si=LCM2E48hVxmO6VG4 https://github.com/Picovoice/picollm PicoLLM is a way to run LLaMa 3 on Node, it will be great for my students. I bet you didn’t know much about Node.js ecosystem for LLMs because it’s very nascent.
Re: The Llama 4 herd
#626Earlier quoted context omitted.
Needing less memory for inference is the entire point of quantization. Saving the disk space or having a smaller download could not justify any level of quality degradation.
Small point of order: > entire point...smaller download could not justify... Q4_K_M has layers and layers of consensus and polling and surveying and A/B testing and benchmarking to show there's ~0 quality degradation. Built over a couple years.
Llama 3.3 already shows a degradation from Q5 to Q4.
As compression improves over the years, the effects of even Q5 quantization will begin to appear
Re: The Llama 4 herd
#627Earlier quoted context omitted.
It is not and has never been half. 2024 voter turnout was 64%
Sure and the voters who did not participate in the election would all have voted the democratic party. I think the election showed that there are real people who apparently don't agree with the democratic party and it would probably be good to listen to these people instead of telling them what to do. (I see the same phenomenon in the Netherlands by the way. The government seems to have decided that they know better…
Re: The Llama 4 herd
#628Earlier quoted context omitted.
Sure and the voters who did not participate in the election would all have voted the democratic party. I think the election showed that there are real people who apparently don't agree with the democratic party and it would probably be good to listen to these people instead of telling them what to do. (I see the same phenomenon in the Netherlands by the way. The government seems to have decided that they know better…
We have an electoral college that essentially disenfranchises any voter that is not voting with the majority unless your state is so close that it could be called a swing state. This affects red state democratic leaning voters just as much as blue state republican leaning voters…their votes are all worthless. For example, the state with the largest number of Trump voters is California, but none of their votes helped…
Re: The Llama 4 herd
#629Earlier quoted context omitted.
Everyone has a "religion" – i.e. a system of values they subscribe to. Secular Americans are annoying because they believe they don't have one, and instead think they're just "good people", calling those who break their core values "bad people".
> Any position is a bias. A flat earther would consider a round-earther biased. That is not what a religion is. > Secular Americans are annoying because they believe they don't have one Why is that a problem to you? > and instead think they're just "good people", calling those who break their core values "bad people". No, not really. Someone is not good or bad because you agree with them. Even a religious person can…
Because seculars/athiests often believe that they're superior to the "stupid, God-believing religious" people, since their beliefs are obviously based on "pure logic and reason".
Yet, when you boil down anyone's value system to its fundamental essence, it turns out to always be a religious-like belief. No human value is based on pure logic, and it's annoying to see someone pretend otherwise.
> Someone is not good or bad because you agree with them
Right, that's what I was arguing against.
> Even a religious person can recognise that an atheist doing charitable work is being good
Sure, but for the sake of argument, I'm honing in on the word "good" here. You can only call something "good" if it aligns with your personal value system.
> The attitude you describe is wrong
You haven't demonstrated how. Could just be a misunderstanding.