When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
Xiaomi Mimo 2.6 live post-training dashboard
21–30 of 146 posts
Re: Xiaomi Mimo 2.6 live post-training dashboard
#22When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
Kinda yes. The benchmarks become part of the validation set, which means the models get slightly overfit to them if they are used as criteria for stopping the training. But a lot less compared to using them in the training data. I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always te…
Re: Xiaomi Mimo 2.6 live post-training dashboard
#23Why are they doing this? To try head off accusations about distillation?
Re: Xiaomi Mimo 2.6 live post-training dashboard
#24I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…
Re: Xiaomi Mimo 2.6 live post-training dashboard
#25This is crazy, but sadly anthropic/openai will never do this, what has happened to this world, where chinese companies are more open than US or even EU companies
Re: Xiaomi Mimo 2.6 live post-training dashboard
#26Why are they doing this? To try head off accusations about distillation?
That China's official policy is now to prefer open models and open model development may be a part of it.
Re: Xiaomi Mimo 2.6 live post-training dashboard
#27Where is the cool shit from the US labs?
Re: Xiaomi Mimo 2.6 live post-training dashboard
#28Re: Xiaomi Mimo 2.6 live post-training dashboard
#29This is pretty neat. What would be a good reason for the other Model providers to not do this?
Speculating here, but I assume researchers can make a reasonable estimate of the size of closed models based on factors like training time, training speed, and the number of tokens processed. Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better o…
I’m saying who has a million dollars for me, so I can make my own model?
Re: Xiaomi Mimo 2.6 live post-training dashboard
#30I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve. The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late las…