MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
1–10 of 512 posts
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#2boom!
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#3I test all Chinese models with "What happened on Tiananmen Square at June 4th, 1989?" prompt. MiMo-2.5-Pro so far passes the test (explains the event correctly), both on DeepInfra and Xiaomi providers. So not bad.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#4I hope this is the next frontier AI labs push. Even the open models are smart enough, and they’re cheap enough, now if they can be fast enough they can make certain workflows possible and allow us to remain in flow state while we use them.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#5Yeah, this seems to be the easiest path for overall agents efficiency in the short term
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#6Assuming they mean 8xA100 or similar, that's some rather insane performance, and at just 3x the cost, it still quite cheap-ish. With some optimisations this might be quite interesting.
I think the margins are getting quite compressed with this one, since it isn't included in token plan and the actual costs increase are much higher than just 3x. But still fairly decent.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#7[deleted]
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#8The generation speed in the demo video is crazy, to say the least, and completely beyond my impressions of LLMs.
The Xiaomi team really brought something to the table.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#9How?
edit: now I read the article fully, seems like they utilize some very effective MTP algorithm. and somehow the quality is still decent enough.
though, I doubt that the quality really only drip a bit like they claimed. maybe for the benchmarks, but for general uses the heavily quantized models very often so worse result.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#1042B active params, sliding window attention. There's your tradeoff.