Earlier quoted context omitted.
You mean all of the frontier models that the Chinese distillation clones are copying? Yeah kinda cool imo. If a dashboard showing training for a model that doesn't even come close to anything us labs have released in 6 months is "cool", then you're a loser
Hahaha. Is that Sam or Dario with throwaway account. This sounds like calling social security, a free handout. Who distills the distillaters? Get it?
Xiaomi Mimo 2.6 live post-training dashboard
131–140 of 146 posts
Re: Xiaomi Mimo 2.6 live post-training dashboard
#132I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…
I am also using 2.5 and it is giving me solid results. Its available free on Openrouter
Re: Xiaomi Mimo 2.6 live post-training dashboard
#133Re: Xiaomi Mimo 2.6 live post-training dashboard
#134I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…
> late last year/early this year That's an eternity when it comes to coding models. In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.
In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good.
I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.
Re: Xiaomi Mimo 2.6 live post-training dashboard
#135Earlier quoted context omitted.
2.6-pro just reached 63.7% by step 10, it's on step 11 right now. Even flash reached 60.7% by step 12, and it's on step 16 now. This is so exciting lmao.
DeepSWE is saturated now IMO, and is basically worthless. Lots of new models get around 74%. Shame too, because it was a pretty decent benchmark for a few months there.
Re: Xiaomi Mimo 2.6 live post-training dashboard
#136Earlier quoted context omitted.
> late last year/early this year That's an eternity when it comes to coding models. In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.
tbf, I the happiest I've been working with claude is late last year/early this year (before March)...
Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.
Now it's *you* being the assistant, reviewer, etc.
Re: Xiaomi Mimo 2.6 live post-training dashboard
#137Earlier quoted context omitted.
> GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better. API may be expensive, but I do 900m tokens (95% cached, ~0.4% output) on Z.ai's $18/mo coding plan with GLM 5.3 Flash .
I wouldn't call that inexpensive. For comparison, I am currently at 6.6B tokens, 95% of monthly quota on a 10$ command code plan, mostly using DeepSeek flash 4.1, or some of the free models for easier tasks.
Re: Xiaomi Mimo 2.6 live post-training dashboard
#138Well, if open source AI is dangerous (for OpenAI/Anthropic IPOs?), this is like watching a time bomb.
For my own usage, Luna is cheap enough that I don't care if other models are cheaper. I'm interested if another model is in some way better and not too expensive.
Re: Xiaomi Mimo 2.6 live post-training dashboard
#139Earlier quoted context omitted.
Not if you don't train against them.
It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used. It's not the direct feedback loop of RL but its not far.
It’s the difference between “study law until you can pass any random bar exam” and “here are 200 legal questions and we’ll drill them, with me correcting and explaining when you get one wrong, until you can pass exactly these 200”.
Your right that tuning can aim for a benchmark, but it does not leak any information about the answers.
Re: Xiaomi Mimo 2.6 live post-training dashboard
#140Earlier quoted context omitted.
Ah, one Donald Trump, a famous neoliberal.
Were Sam Altman and Dario Amodei different men before Trump was in charge?
I don’t think Trump changed them, but Trump is absolutely a symptom of larger social collapse in the US, and that collapse has affected Altman and Amodei. We’re not even pretending that truth matters or that the wealthy can ever suffer consequences, and those two seem quite liberated by that.