Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

441–450 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#442

Am I the only one that doesn’t care about speed? I want it to not do stupid stuff and to be cheaper.

I prefer faster, dumber models because I provide the intelligence myself and I use them only for things that can be verified pretty easily; they do research (with sources) for me, do certain types of code analysis and code search, boilerplate generation, etc., so a fast model is really key.

I don't have any desire (or think it's a good use of LLMs) to one-shot features because even SotA models are incredibly bad at this. I'm optimizing for what they actually seem to be able to do reliably and pretty well, and I want those things to be done fast so I can get on with things.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#443
post #136

So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…

I’m digging into deeper / more complex problems, now. On top of that, I’m also building products faster for our startups, so I am filling in much more of a product role than merely an engineering one. But, really, it is both — and I’m absolutely loving it!

Also, with the added speed I can produce things more in line with the quality I’ve always wanted to add (many more tests, for example).

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#444
post #20

If MiMo v2.5 Pro can run at >1000tk/s on GPUs then I will soon expect the same from OpenAI/Anthropic/Google.

I wouldn't expect any of the american labs to be particularly great (or have much desire) to work on efficiency, they've been consistently proven to be uninterested (if not incapable) of actually improving on those types of things. The closest we've seen lately is that maybe GPT-5.5 (and Opus 4.{7,8}?) are more token-efficient, i.e. they solve things with less tokens...? It hasn't been coupled with any other kind of efficiency bump, though, and we're seeing higher costs anyway in most places where the american labs are involved.

The only players that seem to be capable of a consistent pattern of doing more with less currency are the chinese labs.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#445

i tried to test it and after logging in, i get "You don't have access to this event trial" and can't even log out until i clear my cookies. despite having good model, why such a bad website?

Same. I also found out that my old Xiaomi account is apparently considered "mainland china" and I can't put any phone number except a chinese one on it lol. I'm not trusting these people with anything that's for sure, useless. I'm australian and have never been to china in my life!

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#447
post #337

Earlier quoted context omitted.

> No one is bitter lesson pilled anymore. Will the 10T parameter Mythos model be released this month or next month? They better soon because it is generally accepted that one of the reasons GPT 5.5 is better at hard tasks than Opus is because of its parameter size - and that Opus 4.8 remains competitive only be scaling test-time compute (see how many more tokens it uses than GPT 5.5) https://www.reddit.com/r/LLM/comm…

Why ask me? Anyway, Mythos is not 10T. Anthropic confirmed the training run was under 10^26 flops. You can't train 10T to chincilla and stay under 10^26. Anthropic also confirmed they will not release Mythos, only a "Mythos-class" model, whatever that means.

> Anthropic confirmed the training run was under 10^26 flops. You can't train 10T to chincilla and stay under 10^26.

I don't think Anthropic have said anything of the sort.

Microsoft published it as 6.1*10^27 FLOPs[1]

Elon has claimed the are also training a 10T model because "Some catching up to do"[2]

[1] https://x.com/scaling01/status/2061897540161728791

[2] https://x.com/elonmusk/status/2041754402239975479

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#448

Earlier quoted context omitted.

No. They still have enormous profit margins on inference with these prices.

I highly doubt there is any margin on those inference pricing.

> I highly doubt there is any margin on those inference pricing.

And yet, OpenCode Go offers DeepSeek flash 6 times cheaper than DeepSeek itself. And they claim they are still profitable.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#449
post #346

Earlier quoted context omitted.

I wonder what are the economics driving these pricing decisions? Are the Chinese companies just subsidizing their models to a greater degree than the US, or is this an emergent property of energy policy between countries?

Their models are much smaller: 1T vs 5T for the frontier models. 1T is Sonnet/Google Flash size, not Opus size. The $0.87/M tokens price for Mimo Pro is probably subsidized. Mimo models aren't widely available on western providers, but Kimi and Deepseek are similar sizes and cost about the same to run. They are priced $3-$4/M tokens (which is right were Google's very confused range of Flash models are priced at: betw…

I'm not sure about those parameter sizing claims. Regardless of parameter size, benchmarked intelligence of Chinese and Western frontier models is comparable, so who cares how many parameters it takes to get there.

Mimo is also widely available on western providers. It's on openrouter and you can sign up with Xiaomi directly for a token plan on an English website priced in dollars.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#450
post #335

Earlier quoted context omitted.

It’s near the frontier meaning it’s the best intelligence for the price. It’s not even close to frontier meaning it’s the best intelligence.

I hardly notice DeepSeek being inferior to Claude Opus unless I have it working on tricky and under-defined problems. That is, I trust Opus to reason much better when it has the choice. Otherwise, IME DeepSeek is far cheaper and more effective for anything where the solution is even somewhat obvious.

Out of curiosity, what is your stack? And is this in a legacy project or new one?

I have tried using deep seek flash and pro but they make amateur mistakes. Sonnet level at best.

However v4 flash is absolutely amazing as a generalist model and it’s what we’re using on a product built on top of LLMs. I wish I could code with it but it’s not going to happen anytime soon

Post reply on HN