I'm particularly excited to see a "true base" model to do research off of ( https://huggingface.co/arcee-ai/Trinity-Large-TrueBase ).
I'd love to "chat" to that model see how it behaves
Trinity large: An open 400B sparse MoE model
81–83 of 83 posts
Re: Trinity large: An open 400B sparse MoE model
#82Earlier quoted context omitted.
Progress has not become linear. We've just hit the limits of what we can measure and explain easily. One year ago coding agents could barely do decent auto-complete. Now they can write whole applications. That's much more difficult to show than an ELO score based on how people like emjois and bold text in their chat responses. Don't forget Llama4 led Lmarena and turned out to be very weak.
Much of these gains can be attributed to better tooling and harnesses around the models. Yes, the models also had to be retrained to work with the new tooling, but that doesn’t mean there was a step change in their general “intelligence” or capabilities. And sure enough, I’m seeing the same old flaws as always: frontier models fabricating info not present in the context, having blindness to what is present, getting i…
This isn't the case.
Take Claude Code and use it with Haiku, Sonnet and Opus. There's a huge difference in the capabilities of the models.
> And sure enough, I’m seeing the same old flaws as always: frontier models fabricating info not present in the context, having blindness to what is present, getting into loops, failing to follow simple instructions…
I don't know what frontier models you are using but Opus and Codex 5.2 don't ever do these things for me.
Re: Trinity large: An open 400B sparse MoE model
#83Earlier quoted context omitted.
I'm going to stick to the stuff around Tao, as even well tempered discussion about the rest would be against the guidelines anyways. I had a very different read of Tao's post last month. To me, he opens that there have been many claims of novel solutions which turn out to be known solutions from publications buried for years, but nothing about rapid increase in the rates or even claims mathematicians using LLMs are h…
"...as even well tempered discussion about the rest would be against the guidelines anyways." Didn't bother reading after that. I deeply respect you have the self-awareness to notice and spare us, that's rare. But it also means we all have to have conversations purely on your terms, and because its async, the rules constantly change post-hoc. And that's on top of the post-hoc motte / bailey instances, of which we hav…
If not, continuing to have a conversation can only happen if we want to discuss the recent growth rate of AI and take the time to read what each other write. Similarly, async conversation can be as clear and consistent as we want it to be - we just have to take the time to ask for clarification before writing a response on something we feel could be a movable understanding. Nothing is meant to be unclear as a "gotcha" and I'll always be glad to clarify before moving on.
I also agree nobody should rely solely on LM Arena for benchmarks, which is not what starting a conversation by using it in an example was meant to imply we need to do. I'd love to continue chatting more about other benchmarks and how you see Tao's comments, as you seem to have walked away from reading them with a very different understanding than I did.