Live data from Hacker News

Xiaomi Mimo 2.6 live post-training dashboard

mimo.xiaomi.com

121–130 of 150 posts

Re: Xiaomi Mimo 2.6 live post-training dashboard

#121

For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great. Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort). https://deepswe.datacurve.ai/blog/deepswe-v1-1

gemini 3.8 flash is also 74% and google just started letting all their engineers use claude...go figure

> and google just started letting all their engineers use claude

That's misleading.

1. Having different models available is useful for A/B testing and helping improve Gemini itself.

2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#122
post #106

"Slow down this much openness in AI or we won't get our trillion dollars valuations!" Google had this GPT long go and a wise man within Google noted: "We don't have any maot neither does anyone else." The AI bubble burst is guaranteed and is only delayed by IPOs.

Nothing is guaranteed. Open models have not yet caught up with February's Mythos checkpoint. Meanwhile OpenAI is solving millennium problems, and their compute is still fully utilized.

"Stealing millennium problems" would be more complete if not accurate description. And that 99.99% of the market is not interested in solving millennium problems is the other fact.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#123

Earlier quoted context omitted.

If you're thinking of the UI style, definitely not Claude. It is incapable of writing a clear sentence like "what each step's samples are made of", would have used all-caps for everything, more padding and gradients.

I hope this is /s because it’s very easy to get Claude to write sensibly. That’s why AI slop writing is so annoying because it’s so easy to avoid with any amount of effort at all.

In my experience Opus and Sonnet 5 subtly ignore most instructions related to writing style, and continue to sound the same half of the time. Do you have a successful skill/prompt to share?

Re: Xiaomi Mimo 2.6 live post-training dashboard

#124

For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great. Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort). https://deepswe.datacurve.ai/blog/deepswe-v1-1

2.6-pro just reached 63.7% by step 10, it's on step 11 right now. Even flash reached 60.7% by step 12, and it's on step 16 now. This is so exciting lmao.

DeepSWE is saturated now IMO, and is basically worthless. Lots of new models get around 74%. Shame too, because it was a pretty decent benchmark for a few months there.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#126

I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…

> late last year/early this year That's an eternity when it comes to coding models. In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.

tbf, I the happiest I've been working with claude is late last year/early this year (before March)...

Re: Xiaomi Mimo 2.6 live post-training dashboard

#127
post #106

"Slow down this much openness in AI or we won't get our trillion dollars valuations!" Google had this GPT long go and a wise man within Google noted: "We don't have any maot neither does anyone else." The AI bubble burst is guaranteed and is only delayed by IPOs.

I wonder what we will do with the discarded data centers and its hardware..

Copper can be stole, but that's already in progress.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#128

Earlier quoted context omitted.

You mean all of the frontier models that the Chinese distillation clones are copying? Yeah kinda cool imo. If a dashboard showing training for a model that doesn't even come close to anything us labs have released in 6 months is "cool", then you're a loser

Why the fuck would you or I care about that? Anthropic and OpenAI literally stole from every human in history and youre out here complaining that the Chinese are distilling models and releasing them to the public? Why do you care?

Because without those labs to distill from the pathetic Chinese labs wouldn't have anything. Im not impressed by them copying US labs not sure why you are. But go off ccp bot

Re: Xiaomi Mimo 2.6 live post-training dashboard

#129
post #101

Earlier quoted context omitted.

Haha yeah pretty wild how easily you can see the data is fake by the repeating numbers (refresh the page the progress goes back in time constantly) + watch for restarts. They say they happen but 0 data correlates the log messages. Just a replay of old data or being fed by an llm so they convince people they are open

The intermediate tickers are fake but real data comes in and resets it. Its like a progress bar essentially. We don't call progress and bars fake

I do when their fake like this site is. Insane people blindly believe this stuff

Re: Xiaomi Mimo 2.6 live post-training dashboard

#130

This is so very clearly fake? See the message stating the flash 2.6 flash run was restarted and 0 graphs correlate that restart

A restart of the process does not necessarily mean reverting the model state. I don't know why you would even do that, because you'd lose all the progress you made.

It said restarted step 15 5 mins ago and the progress showed they were working on step 16 for a day
Post reply on HN