Live data from Hacker News

Xiaomi Mimo 2.6 live post-training dashboard

mimo.xiaomi.com

41–50 of 146 posts

Re: Xiaomi Mimo 2.6 live post-training dashboard

#41
post #33

The Chinese labs are just making fun of the US labs at this point. Where is the cool shit from the US labs?

With other software, devs convince their managers of the importance of using open source stuff in their stack. With AI, it's usually managers choosing what models to use for the devs. The US labs don't need to give a damn how much devs like open source

> The US labs don't need to give a damn how much devs like open source

In the short term, true.

In the long term, unknown but typically when you hold progress that way while other countries don't you at best end up becoming siloed while the rest of the world continues on without you.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#42
post #7

When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.

They exist to detect degradation. Datasets are not perfect and if a batch contains too much bad data it can ruin a run, also an opportunity to find bad data and improve the dataset filtering.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#43
post #39
post #7

When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.

Not if you don't train against them.

It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.

It's not the direct feedback loop of RL but its not far.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#45
post #33

The Chinese labs are just making fun of the US labs at this point. Where is the cool shit from the US labs?

With other software, devs convince their managers of the importance of using open source stuff in their stack. With AI, it's usually managers choosing what models to use for the devs. The US labs don't need to give a damn how much devs like open source

This isn't about liking open source. This is about the labs just being cool and doing cool shit instead of the opposite which is Anthropic where all they talking about is killing everyone and taking everyone's job.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#46
Neat! I've been trying out their next model for the last week, which I assume is a version of this, and it's been a good experience so far.

I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.

The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#48

I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…

I've found that mimo v2.5 works for very basic things like a python script to do one thing, but it also is very 'dumb' compared to qwen 3.8-flash-next (I think the benchmark scores for terminal and coding specific benches back this up). And definitely not in the same class as like a GLM5.2 or 5.3. It's fast but makes basic mistakes that only get caught later.

[deleted]

Re: Xiaomi Mimo 2.6 live post-training dashboard

#49
post #38

I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…

How fast is it compared with the other Chinese models?

They both are in the 50-100 tok/s range. The Mimo v2.5 Pro Ultraspeed beta could reach 1000 tok/s, hoping they can do something similar for the new model, it was amazing.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#50

You'd think they would make it less obvious that they are running their whole operation with Claude

If you're thinking of the UI style, definitely not Claude. It is incapable of writing a clear sentence like "what each step's samples are made of", would have used all-caps for everything, more padding and gradients.
Post reply on HN