MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
441–450 of 512 posts
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#442Am I the only one that doesn’t care about speed? I want it to not do stupid stuff and to be cheaper.
I don't have any desire (or think it's a good use of LLMs) to one-shot features because even SotA models are incredibly bad at this. I'm optimizing for what they actually seem to be able to do reliably and pretty well, and I want those things to be done fast so I can get on with things.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#443So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…
Also, with the added speed I can produce things more in line with the quality I’ve always wanted to add (many more tests, for example).
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#444If MiMo v2.5 Pro can run at >1000tk/s on GPUs then I will soon expect the same from OpenAI/Anthropic/Google.
The only players that seem to be capable of a consistent pattern of doing more with less currency are the chinese labs.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#445i tried to test it and after logging in, i get "You don't have access to this event trial" and can't even log out until i clear my cookies. despite having good model, why such a bad website?
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#446Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#447Earlier quoted context omitted.
> No one is bitter lesson pilled anymore. Will the 10T parameter Mythos model be released this month or next month? They better soon because it is generally accepted that one of the reasons GPT 5.5 is better at hard tasks than Opus is because of its parameter size - and that Opus 4.8 remains competitive only be scaling test-time compute (see how many more tokens it uses than GPT 5.5) https://www.reddit.com/r/LLM/comm…
Why ask me? Anyway, Mythos is not 10T. Anthropic confirmed the training run was under 10^26 flops. You can't train 10T to chincilla and stay under 10^26. Anthropic also confirmed they will not release Mythos, only a "Mythos-class" model, whatever that means.
I don't think Anthropic have said anything of the sort.
Microsoft published it as 6.1*10^27 FLOPs[1]
Elon has claimed the are also training a 10T model because "Some catching up to do"[2]
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#448Earlier quoted context omitted.
No. They still have enormous profit margins on inference with these prices.
I highly doubt there is any margin on those inference pricing.
And yet, OpenCode Go offers DeepSeek flash 6 times cheaper than DeepSeek itself. And they claim they are still profitable.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#449Earlier quoted context omitted.
I wonder what are the economics driving these pricing decisions? Are the Chinese companies just subsidizing their models to a greater degree than the US, or is this an emergent property of energy policy between countries?
Their models are much smaller: 1T vs 5T for the frontier models. 1T is Sonnet/Google Flash size, not Opus size. The $0.87/M tokens price for Mimo Pro is probably subsidized. Mimo models aren't widely available on western providers, but Kimi and Deepseek are similar sizes and cost about the same to run. They are priced $3-$4/M tokens (which is right were Google's very confused range of Flash models are priced at: betw…
Mimo is also widely available on western providers. It's on openrouter and you can sign up with Xiaomi directly for a token plan on an English website priced in dollars.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#450Earlier quoted context omitted.
It’s near the frontier meaning it’s the best intelligence for the price. It’s not even close to frontier meaning it’s the best intelligence.
I hardly notice DeepSeek being inferior to Claude Opus unless I have it working on tricky and under-defined problems. That is, I trust Opus to reason much better when it has the choice. Otherwise, IME DeepSeek is far cheaper and more effective for anything where the solution is even somewhat obvious.
I have tried using deep seek flash and pro but they make amateur mistakes. Sonnet level at best.
However v4 flash is absolutely amazing as a generalist model and it’s what we’re using on a product built on top of LLMs. I wish I could code with it but it’s not going to happen anytime soon