MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
11–20 of 512 posts
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#12Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#13Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#1442B active params, sliding window attention. There's your tradeoff.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#15I test all Chinese models with "What happened on Tiananmen Square at June 4th, 1989?" prompt. MiMo-2.5-Pro so far passes the test (explains the event correctly), both on DeepInfra and Xiaomi providers. So not bad.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#16Discussions about choosing a library with the best syntactic sugar method naming is just as crazy as suggesting we type in assembly.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#17Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#18Assuming they mean 8xA100 or similar, that's some rather insane performance, and at just 3x the cost, it still quite cheap-ish. With some optimisations this might be quite interesting. I think the margins are getting quite compressed with this one, since it isn't included in token plan and the actual costs increase are much higher than just 3x. But still fairly decent.
Remember, these guys are not VC backed. Anything they do must break even