Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

451–460 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#451
post #180

Earlier quoted context omitted.

I'm kind of poor so I have been trying to use DeepSeek v4 Flash, GLM 5.1 etc. as much as possible recently instead of Claude or GPT.

You would do us all a service by telling us how your experiences of that have been.

The only one that is really close to Claude in performance is GLM-5.1. The others (Mimo, deepseek, etc..) looks good on paper but usually fails on a multi-step agentic orchestration.

This is at least my experience with Claude Code as harness. Also, GLM pricing is not that far off from Claude. It's cheaper but not DeepSeek cheap.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#452
post #438

Earlier quoted context omitted.

I'm pretty sure Xi is also a sociopath, but he differs from Trump in that he's competent. And maybe that's a good thing for American democracy--if we had a competent dictator who could manifest massive infrastructure projects maybe the pro-democracy backlash would be significantly attenuated?

Oh, I was thinking of OpenAI and Anthropic CEOs.

Heh, isn’t it fun living in a timeline where there are so many sociopathic leaders that your earlier comment is ambiguous? (:

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#453
post #156
post #19

Cerebras is trialing Kimi K2.6 at 3000t/s (invite only). I'm excited for when the fast hardware gets more mainstream for frontier models. Models designed for speed on Nvidia are nice addition that could bridge the gap.

Source? Their website says 1000t/s https://www.cerebras.ai/blog/which-is-faster-gemini-3-5-flas...

This is likely correct, sorry for the bad info. Was working from memory.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#454
post #99

I don't understand, given all they say, why this would not be made available to everyone at once? Why the limited release? They should have no trouble scaling it if it runs on a single rack.

I wonder about this too. The other objections miss the point: if it's faster, and otherwise the same, and doesn't require different hardware, then why not just announce that the standard tier of MiMo-v.25-Pro is now ridiculously fast and raise the price? What does "limited high speed resources" mean if it runs on the same hardware as the rest of their pool? I think the answer is that there's a tradeoff here where add…

> and doesn't require different hardware

But it may well do. They mention TileRT in the announcement, so this speed comes from low level optimization for some specific GPU target.

With availability of SOTA western GPUs being scarce in China, they may well have a mishmash of different GPUs.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#455

Cool, what is the price pr. Million token. I am using a 300 t/s model for a project I am doing and speed is crucial over precision, so this seems like an upgrade. However if it is 10$ pr. M tokens then it is not worth an upgrade.

$0.435/$0.87 for the standard speed, this one should be 3 times that.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#456
post #99

Earlier quoted context omitted.

I wonder about this too. The other objections miss the point: if it's faster, and otherwise the same, and doesn't require different hardware, then why not just announce that the standard tier of MiMo-v.25-Pro is now ridiculously fast and raise the price? What does "limited high speed resources" mean if it runs on the same hardware as the rest of their pool? I think the answer is that there's a tradeoff here where add…

> and doesn't require different hardware But it may well do. They mention TileRT in the announcement, so this speed comes from low level optimization for some specific GPU target. With availability of SOTA western GPUs being scarce in China, they may well have a mishmash of different GPUs.

They specifically said it's stock hardware, but... yeah, maybe highly specific stock hardware.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#457
post #58

Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…

I'm using Deepseek-v4-pro as my main model and this is sometimes pretty annoying, I have to do some easy boring task, think "I'll just leave the agent to do it and go take a nap", but it's already done writing the code before I even walk away from the computer

Take the nap anyway, just say it took all afternoon :)

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#458

Earlier quoted context omitted.

No. They still have enormous profit margins on inference with these prices.

Any source to backup this claim, pretty please?

Source? There are a countless number of providers serving open weight models for fun and profit.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#459
post #406

Earlier quoted context omitted.

Have you tried https://chatjimmy.ai/ it’s only a demo but it blew my mind. I had the sudden feeling that this is the future.

What do you mean "demo"? Seems to work... Who is behind this?

These guys: https://taalas.com/products/

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#460
post #447

Earlier quoted context omitted.

Why ask me? Anyway, Mythos is not 10T. Anthropic confirmed the training run was under 10^26 flops. You can't train 10T to chincilla and stay under 10^26. Anthropic also confirmed they will not release Mythos, only a "Mythos-class" model, whatever that means.

> Anthropic confirmed the training run was under 10^26 flops. You can't train 10T to chincilla and stay under 10^26. I don't think Anthropic have said anything of the sort. Microsoft published it as 6.1*10^27 FLOPs[1] Elon has claimed the are also training a 10T model because "Some catching up to do"[2] [1] https://x.com/scaling01/status/2061897540161728791 [2] https://x.com/elonmusk/status/2041754402239975479

I must have confused mythos with opus 4.7. One of their recent model cards confirmed that training flops was under the EO reporting requirement of 10^26 flops.
Post reply on HN