Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

81–90 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#81
post #58

Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…

We fit in for the things that are not artificial.

So long as AI lives in server farms, humans will be needed for tasks in the physical world.

It's only if we combine AI with robots that things get really dicey.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#82

Earlier quoted context omitted.

Suspect this will be included once out of beta but at a higher credit/token ratio. Remember, these guys are not VC backed. Anything they do must break even

Chinese "companies" are not companies in the western sense, but more like government departments with capitalist styling to deceive the western audience. From that point of view, they have as much money as they need. That's why there is no "VC", because Chinese government assumes that role.

Huge L for free market economies if true

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#84
post #6

Assuming they mean 8xA100 or similar, that's some rather insane performance, and at just 3x the cost, it still quite cheap-ish. With some optimisations this might be quite interesting. I think the margins are getting quite compressed with this one, since it isn't included in token plan and the actual costs increase are much higher than just 3x. But still fairly decent.

Must be Blackwell for native fp4 support.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#85
post #19

Cerebras is trialing Kimi K2.6 at 3000t/s (invite only). I'm excited for when the fast hardware gets more mainstream for frontier models. Models designed for speed on Nvidia are nice addition that could bridge the gap.

now that's what i call a software development breakthrough/platform! thanks for the heads up!

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#86

Given that MiMo is as cheap as Deepseek ( previous discussion: https://news.ycombinator.com/item?id=48282814 ) multiplying that by 3x for ultra speed is still shockingly cheap.

MiMo and DeepSeek are not cheap. Anthropic and OpenAI are expensive for what they provide.

You don't consider Input $0.435 Output $0.87 cache read $0.003625 per million tokens for near frontier intelligence cheap?

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#87
post #3

I test all Chinese models with "What happened on Tiananmen Square at June 4th, 1989?" prompt. MiMo-2.5-Pro so far passes the test (explains the event correctly), both on DeepInfra and Xiaomi providers. So not bad.

What would be a correct explanation of the event?

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#88
post #16

I may sound like a shill, but exponential growth and all. We are going to get near instant software from prompt, multiple ones and then choose the best one. Discussions about choosing a library with the best syntactic sugar method naming is just as crazy as suggesting we type in assembly.

And how are you going to determine which is the best? Going through all the possible combinations of users and usage? So mostly it shifts the work from generation to validation.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#89
post #15

Earlier quoted context omitted.

No idea why you've been downvoted. This is excellent news.

Because this never gets brought up about US models, which have just as much censorship as the Chinese ones.

No, US models have alignment. Only Chinese models have censorship.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#90
post #9

How? edit: now I read the article fully, seems like they utilize some very effective MTP algorithm. and somehow the quality is still decent enough. though, I doubt that the quality really only drip a bit like they claimed. maybe for the benchmarks, but for general uses the heavily quantized models very often so worse result.

They say they are using https://github.com/tile-ai/TileRT

- persistent CUDA kernel

- tiled processing with overlapping read/writes

- model designed with specific constraints in mind

Post reply on HN