Live data from Hacker News

Trinity large: An open 400B sparse MoE model

arcee.ai

81–83 of 83 posts

Re: Trinity large: An open 400B sparse MoE model

#81

I'm particularly excited to see a "true base" model to do research off of ( https://huggingface.co/arcee-ai/Trinity-Large-TrueBase ).

I'd love to "chat" to that model see how it behaves

I've done this out of curiosity with the base model of LLama 3.1 405B. I vibe coded a little chat harness with the system prompt being a few short conversations between "system" and "user" with "user:" being the stop word so I could enter my message. Worked surprisingly well and I didn't get any sycophancy or cliched AI responses.

Re: Trinity large: An open 400B sparse MoE model

#82
post #38

Earlier quoted context omitted.

Progress has not become linear. We've just hit the limits of what we can measure and explain easily. One year ago coding agents could barely do decent auto-complete. Now they can write whole applications. That's much more difficult to show than an ELO score based on how people like emjois and bold text in their chat responses. Don't forget Llama4 led Lmarena and turned out to be very weak.

Much of these gains can be attributed to better tooling and harnesses around the models. Yes, the models also had to be retrained to work with the new tooling, but that doesn’t mean there was a step change in their general “intelligence” or capabilities. And sure enough, I’m seeing the same old flaws as always: frontier models fabricating info not present in the context, having blindness to what is present, getting i…

> Much of these gains can be attributed to better tooling and harnesses around the models.

This isn't the case.

Take Claude Code and use it with Haiku, Sonnet and Opus. There's a huge difference in the capabilities of the models.

> And sure enough, I’m seeing the same old flaws as always: frontier models fabricating info not present in the context, having blindness to what is present, getting into loops, failing to follow simple instructions…

I don't know what frontier models you are using but Opus and Codex 5.2 don't ever do these things for me.

Re: Trinity large: An open 400B sparse MoE model

#83

Earlier quoted context omitted.

I'm going to stick to the stuff around Tao, as even well tempered discussion about the rest would be against the guidelines anyways. I had a very different read of Tao's post last month. To me, he opens that there have been many claims of novel solutions which turn out to be known solutions from publications buried for years, but nothing about rapid increase in the rates or even claims mathematicians using LLMs are h…

"...as even well tempered discussion about the rest would be against the guidelines anyways." Didn't bother reading after that. I deeply respect you have the self-awareness to notice and spare us, that's rare. But it also means we all have to have conversations purely on your terms, and because its async, the rules constantly change post-hoc. And that's on top of the post-hoc motte / bailey instances, of which we hav…

The conversation is certainly not on "my terms" as I didn't write the guidelines (nor do they benefit me more than anyone else). If you are genuinely concerned with the conversation, please flag it and/or email hn@ycombinator.com and they will (genuinely) handle it appropriately. Otherwise there is not much else which can be said around this here.

If not, continuing to have a conversation can only happen if we want to discuss the recent growth rate of AI and take the time to read what each other write. Similarly, async conversation can be as clear and consistent as we want it to be - we just have to take the time to ask for clarification before writing a response on something we feel could be a movable understanding. Nothing is meant to be unclear as a "gotcha" and I'll always be glad to clarify before moving on.

I also agree nobody should rely solely on LM Arena for benchmarks, which is not what starting a conversation by using it in an example was meant to imply we need to do. I'd love to continue chatting more about other benchmarks and how you see Tao's comments, as you seem to have walked away from reading them with a very different understanding than I did.

Post reply on HN