Earlier quoted context omitted.
What Chinese firms are doing makes perfect sense from the commercial perspective actually because they understand how a classic commoditization spiral works. The reality is that models themselves are general commodities and there's just not enough difference between them. A company can get ahead of others by a few months, but then the rest quickly close the gap. It's a really low margin business because there's no wa…
The Chinese labs incentives is to run inference for the world, because inference can run on the homegrown Chinese chips (giving them a guaranteed market for their hardware) and they have cheap plentiful power. The US frontier labs have an incentive to do deals with large firms to act like a contract research organization, taking royalties on creations/discoveries. Alex Karp called this out in his rant ("Why charge fo…
Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
51–53 of 53 posts
I expect that US companies will ultimately angle for long term government contracts similarly to companies like Raytheon. There's basically unlimited money available, and they don't have to worry about competition from China here.
Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#52Earlier quoted context omitted.
There is no such mandate. ByteDance keeps their models closed. So does iFlyTek. Qwen Max is closed as well.
But they are not groundbreaking. They are simply copies of other similar Chinese models The mandate is worded differently from what I said
What is the literal wording of this "mandate" then?
Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#53Related, I was given access to mimo-v2.5-ultraspeed, which is amazing. This is now my expectation for speed, it’s fast enough for me to stay mentally engaged rather than getting distracted waiting for the agent to churn.
Is it the same quality as base mimo v2.5, or different? I've been enjoying regular mimo v2.5 quite a bit via opencode, if ultraspeed provides the same quality, that's crazy.
I'm not a regular v2.5 user, so I don't really know. But given the TileRT team write-up says the entire network gets quantized to FP8 (and experts get quantized to FP4)[0], I'm assuming there's at least a modest drop in quality.