Live data from Hacker News

Xiaomi Mimo 2.6 live post-training dashboard

mimo.xiaomi.com

91–100 of 151 posts

Re: Xiaomi Mimo 2.6 live post-training dashboard

#91
post #73

Earlier quoted context omitted.

The fact I can run Qwen 3.8 Flash Next locally, forever (on my DGX Spark-alike) is genuinely shocking to me. It’s crazy good for how small it is. Fast, too.

Yeah, I'm guessing you have a variant that fits in <128GB with 262k context? I have the unsloth Q8 GGUF of it here in a setup that with full context and ton of extra llama-server "--cache-ram" sits around 200GB RAM usage on a 256GB system, it's probably the best thing I've found for a 256GB class machine. Enough headroom for a rope/yarn extension to 524288 context if I need it.

Yep, the engrams are on NVMe (the speed penalty was lower than I expected) and it is quantised to fit.

It’s good enough that I’m considering a second spark, or selling this and buying an M5 Ultra with 256GB for it

Re: Xiaomi Mimo 2.6 live post-training dashboard

#92

This is so very clearly fake? See the message stating the flash 2.6 flash run was restarted and 0 graphs correlate that restart

A restart of the process does not necessarily mean reverting the model state. I don't know why you would even do that, because you'd lose all the progress you made.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#94
post #73

Earlier quoted context omitted.

I've found that mimo v2.5 works for very basic things like a python script to do one thing, but it also is very 'dumb' compared to qwen 3.8-flash-next (I think the benchmark scores for terminal and coding specific benches back this up). And definitely not in the same class as like a GLM5.2 or 5.3. It's fast but makes basic mistakes that only get caught later.

The fact I can run Qwen 3.8 Flash Next locally, forever (on my DGX Spark-alike) is genuinely shocking to me. It’s crazy good for how small it is. Fast, too.

I made this 3D game in a day on the same setup with Qwen Code as agent: https://games.jonathanpage.com/

And I am not a web developer! It's an extraordinary model.

(Mouse and keyboard required)

Re: Xiaomi Mimo 2.6 live post-training dashboard

#95

The Chinese labs are just making fun of the US labs at this point. Where is the cool shit from the US labs?

You mean all of the frontier models that the Chinese distillation clones are copying? Yeah kinda cool imo. If a dashboard showing training for a model that doesn't even come close to anything us labs have released in 6 months is "cool", then you're a loser

Why the fuck would you or I care about that?

Anthropic and OpenAI literally stole from every human in history and youre out here complaining that the Chinese are distilling models and releasing them to the public?

Why do you care?

Re: Xiaomi Mimo 2.6 live post-training dashboard

#96

This is crazy, but sadly anthropic/openai will never do this, what has happened to this world, where chinese companies are more open than US or even EU companies

Neoliberalism, that famously open and transparent economic ideology

Ah, one Donald Trump, a famous neoliberal.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#97

The Chinese labs are just making fun of the US labs at this point. Where is the cool shit from the US labs?

You mean all of the frontier models that the Chinese distillation clones are copying? Yeah kinda cool imo. If a dashboard showing training for a model that doesn't even come close to anything us labs have released in 6 months is "cool", then you're a loser

No crying in the copyright casino.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#99
post #12
post #3

This is pretty neat. What would be a good reason for the other Model providers to not do this?

Speculating here, but I assume researchers can make a reasonable estimate of the size of closed models based on factors like training time, training speed, and the number of tokens processed. Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better o…

I don't really remember a situation, which of those models supposedly beat the other?

I still opus 4.6 though not for code

Re: Xiaomi Mimo 2.6 live post-training dashboard

#100

I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…

Same here. It’s the first AI provider I actually gave money to, since they offered the model for free with a Mimo code for the first month or so, and it was great.

These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.

Post reply on HN