Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
1–10 of 53 posts
Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#2It does inspire hope that the Chinese labs seem to be so open although the sceptic in me does wonder what their end game is.
Surely, from a purely economic perspective it would be wiser to keep this proprietary and benefit from the increased API traffic?
Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#3I've used Mimo extensively in the past few months, can't wait to see what v3 will bring.
Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#4It's really cool and interesting to see the kind of engineering that goes into Xiaomi (and Deepseeks) inference optimizations. Z.ai has also published some interesting papers although I haven't had a chance to go through them yet. It does inspire hope that the Chinese labs seem to be so open although the sceptic in me does wonder what their end game is. Surely, from a purely economic perspective it would be wiser to…
At Xiaomi, MiMo is now led by Luo Fuli. She is a former Alibaba & DeepSeek employee: https://newsen.pku.edu.cn/news_events/news/people/15385.html (https://archive.vn/I8Pmu) / https://e.vnexpress.net/news/tech/personalities/who-is-luo-f... (https://archive.vn/sb3B6)
Don't know if it is due to Luo, but it is striking how similar performance & pricing of the models, DeepSeek v4 Pro & MiMo v2.5 Pro, is.
Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#5Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#6Will they lower the price or is this documenting past work
Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#7Will they lower the price or is this documenting past work
Their pricing is incredible on the token plan - something like 50b tokens for $60!
Someone did the math a few months ago and paying API prices was the same as the monthly subscription.
Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#8It's really cool and interesting to see the kind of engineering that goes into Xiaomi (and Deepseeks) inference optimizations. Z.ai has also published some interesting papers although I haven't had a chance to go through them yet. It does inspire hope that the Chinese labs seem to be so open although the sceptic in me does wonder what their end game is. Surely, from a purely economic perspective it would be wiser to…
Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#9It's really cool and interesting to see the kind of engineering that goes into Xiaomi (and Deepseeks) inference optimizations. Z.ai has also published some interesting papers although I haven't had a chance to go through them yet. It does inspire hope that the Chinese labs seem to be so open although the sceptic in me does wonder what their end game is. Surely, from a purely economic perspective it would be wiser to…
Re: Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
#10It's really cool and interesting to see the kind of engineering that goes into Xiaomi (and Deepseeks) inference optimizations. Z.ai has also published some interesting papers although I haven't had a chance to go through them yet. It does inspire hope that the Chinese labs seem to be so open although the sceptic in me does wonder what their end game is. Surely, from a purely economic perspective it would be wiser to…
Publishing open weights gives me more confidence in the model, and ironically makes me less anxious about making sure I can replace the cloud usage with a local alternative. Whereas I’m very nervous right now with relying on 5.6-Sol - what if they triple the price, nerf it, etc.?