Live data from Hacker News

Xiaomi Mimo 2.6 live post-training dashboard

mimo.xiaomi.com

111–120 of 151 posts

Re: Xiaomi Mimo 2.6 live post-training dashboard

#113
post #106

"Slow down this much openness in AI or we won't get our trillion dollars valuations!" Google had this GPT long go and a wise man within Google noted: "We don't have any maot neither does anyone else." The AI bubble burst is guaranteed and is only delayed by IPOs.

I wonder what we will do with the discarded data centers and its hardware..

Re: Xiaomi Mimo 2.6 live post-training dashboard

#114
post #106

"Slow down this much openness in AI or we won't get our trillion dollars valuations!" Google had this GPT long go and a wise man within Google noted: "We don't have any maot neither does anyone else." The AI bubble burst is guaranteed and is only delayed by IPOs.

Nothing is guaranteed.

Open models have not yet caught up with February's Mythos checkpoint.

Meanwhile OpenAI is solving millennium problems, and their compute is still fully utilized.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#115

For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great. Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort). https://deepswe.datacurve.ai/blog/deepswe-v1-1

gemini 3.8 flash is also 74% and google just started letting all their engineers use claude...go figure

Re: Xiaomi Mimo 2.6 live post-training dashboard

#116

I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…

> late last year/early this year

That's an eternity when it comes to coding models.

In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#119

Earlier quoted context omitted.

this is the rl run, not the pretraining run

even in pre-training, usually 30%-50% is code these days.

That would be far too high in my opinion. But happy if anybody can give insights from their own experience with pretraining runs.
Post reply on HN