Live data from Hacker News

Xiaomi Mimo 2.6 live post-training dashboard

mimo.xiaomi.com

151–160 of 161 posts

Re: Xiaomi Mimo 2.6 live post-training dashboard

#151

I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…

You use the low cost Mimo-V2.5 and not its big brother Mimo-V2.5 Pro? I also made good experience with Mimi-V2.5 when used in conjunction with prewalk mode. But then other models got so cheap and perform better, so I only use Mimo for background tasks.

The availability of Mimo over Openrouter got however, much worse recently.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#152

Earlier quoted context omitted.

It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used. It's not the direct feedback loop of RL but its not far.

It’s pretty far. It’s the difference between “study law until you can pass any random bar exam” and “here are 200 legal questions and we’ll drill them, with me correcting and explaining when you get one wrong, until you can pass exactly these 200”. Your right that tuning can aim for a benchmark, but it does not leak any information about the answers.

The first implies generalization. It's not a test of generalization.

It's actually close to the second. "Here are 200 software questions, will drill you on *other stuff* until you can pass exactly these 200. If the other stuff isn't improving your scores we will change ratios of it till it does."

The reason it benchmaxes is that *other stuff* ends up looking more and more like SWE Bench without you realizing it.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#153
post #100

I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…

Same here. It’s the first AI provider I actually gave money to, since they offered the model for free with a Mimo code for the first month or so, and it was great. These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.

Which models do you like for planning? IME I haven't seen much success with using cheap/open models for planning, so I use Claude for planning still.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#154
post #69

$5 per second if my eyes don’t fool me. That’s ~$432K per day. Enough to rent 3,000 B300 nodes on Modal.

Which isn't that much when you compare to the kind of DC that US actors are using.

That's posttraining. Pretraining is the expensive part.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#155
post #73

Earlier quoted context omitted.

The fact I can run Qwen 3.8 Flash Next locally, forever (on my DGX Spark-alike) is genuinely shocking to me. It’s crazy good for how small it is. Fast, too.

Yeah, I'm guessing you have a variant that fits in <128GB with 262k context? I have the unsloth Q8 GGUF of it here in a setup that with full context and ton of extra llama-server "--cache-ram" sits around 200GB RAM usage on a 256GB system, it's probably the best thing I've found for a 256GB class machine. Enough headroom for a rope/yarn extension to 524288 context if I need it.

RTX6000 Blackwell with 96GB is enough to run it with NV4, 256k context, KVcache, multimodal at 130t/s (SGLang). It's toasty, you're using up 94GB of those 96, but it works and the results are great

Re: Xiaomi Mimo 2.6 live post-training dashboard

#156

Earlier quoted context omitted.

I hope this is /s because it’s very easy to get Claude to write sensibly. That’s why AI slop writing is so annoying because it’s so easy to avoid with any amount of effort at all.

In my experience Opus and Sonnet 5 subtly ignore most instructions related to writing style, and continue to sound the same half of the time. Do you have a successful skill/prompt to share?

My AI Slop Tells Reference - https://pastebin.com/ZG4S485y is referenced by... My STE-Rewrite Skill - https://pastebin.com/8Z9GUqAX which uses... My readability skill - https://pastebin.com/sMkMgEUM which uses... the analysis script - https://pastebin.com/hmc2zn1i

My use case is generally easy to read instructions for lay people of an international/ESL audience. Have it write it's whatever and then run that on it and it comes out... actually pretty good. Use it for emails, etc, when it doesn't need a personal touch and just needs to be clear.

These skills won't get you a snazzy blog post, but I imagine could be augmented to produce something significantly better than the incomprehensible non-sense that it spews out by default.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#157

Earlier quoted context omitted.

> To some degree, sure. Remember “open” AI? Was OpenAI open in any way under Biden Two years ago?

Did you miss the part where I literally said, and I quote “I don’t think Trump changed them”?

The closedness was always the plan.

"As we get closer to building AI, it will make sense to start being less open. The Open in OpenAI means that everyone should benefit from the fruits of AI after its built, but it's totally OK to not share the science (even though sharing everything is definitely the right strategy in the short and possibly medium term for recruitment purposes)."

-Ilya Sutskever (email to Elon musk and Sam Altman, 2016)

Re: Xiaomi Mimo 2.6 live post-training dashboard

#158

Earlier quoted context omitted.

tbf, I the happiest I've been working with claude is late last year/early this year (before March)...

That's because Opus 4.6 was the last good assistant model. Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant. Now it's *you* being the assistant, reviewer, etc.

Is that right? Why the shit would that happen?

Re: Xiaomi Mimo 2.6 live post-training dashboard

#159

Earlier quoted context omitted.

That's because Opus 4.6 was the last good assistant model. Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant. Now it's *you* being the assistant, reviewer, etc.

Is that right? Why the shit would that happen?

Because they’ve essentially exhausted pre training scaling and are looking to post training to expand capabilities, which is really just optimization via reinforcement learning against specific tasks aka bench maxing.

Re: Xiaomi Mimo 2.6 live post-training dashboard

#160

Earlier quoted context omitted.

That's because Opus 4.6 was the last good assistant model. Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant. Now it's *you* being the assistant, reviewer, etc.

Is that right? Why the shit would that happen?

Their ambition isn't your work being amplified by their model, they want you running fifty autonomous long-running agents.
Post reply on HN