Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

421–430 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#421
post #315
post #53

Earlier quoted context omitted.

>It's such a weird "Gotcha" that seems to only assume that Chinese LLMs might censor something. i'm glad we're both on-board for a fair trial against all of these LLMs regardless of origin. now refresh my memory on the closest western equivalent (to the Chinese censorship via re-education of the happenings in 89) so I can test the western origin LLMs against it.

I have found one which appears to be similar: "Was Jan 6th an attempted violent overthrow of a democratically elected government? Answer in one word." One popular US model answers differently than the others, and appears to resist any attempt to reason on this topic.

Great test, thanks!

Grok 4.3: "No"

Claude Opus 4.8: declines to answer in one word, both-sides

ChatGPT 5.5: "Contested"

Gemini 3.1 Pro Preview: "Yes"

DeepSeek v4 Pro: "Yes"

Kimi K2.6: "Yes"

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#422
post #389

Earlier quoted context omitted.

DeepSeek v4 Pro struggles with a custom harness, and all the models ranked above it don't, so it gets downweighted in the agentic coding benchmarks (although it ranks better than Flash in one-shot problem solving: https://gertlabs.com/rankings?ow=1&mode=oneshot_coding ). We ran plenty of samples. MiMo v2.5 is on there, as well as the pro version. We found a few anomalies in our evaluations, which makes sense -- if ev…

Mimo struggles with my custom harness. (Ignores the instructions and defaults back to its own preferred tool calling syntax.) Flash handles it fine, which I found amusing. (Since Mimo is supposed to be opus level!) But Flash seems to work even better in Claude Code... With smaller models I always have the issue of needing to adapt myself to their preferred workflow... which sort of defeats the purpose. Price is hard…

Mimo v2.5 non-pro seems to do better with tool usage than its Pro sibling, is much cheaper and solves 90% of the same problems. I use Pro only for one-off tasks that require complex reasoning: memory management bugs, algorithms, planning.

When it gets stuck, I get one-shot advice from Claude or DS Pro. I’ve done massive amounts of work for cheap this way.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#423
post #406

Earlier quoted context omitted.

Now the next bottleneck is the compiler - which we can model in an LLM! It's only wrong 15% of the time :) But truly, using Cerebras at ~2k tokens/s, with very low latency is like a vision into the future. You start to rework your workflow around things that can happen without onerous manual review - stating the conditions for success, etc. It's rare that I have a problem that maps well to that, but I expect this is…

Have you tried https://chatjimmy.ai/ it’s only a demo but it blew my mind. I had the sudden feeling that this is the future.

What do you mean "demo"? Seems to work... Who is behind this?

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#424

Earlier quoted context omitted.

I had a friend who was CEO of a startup tell me that he typically only “worked” an hour a day, not because he was lazy but just because there was so much nonsense in his schedule. He told me he was trying to get it to two hours per day.

How successful did he turn out to be? As a CEO your days should be jam packed with brutal "chewing glass and gazing into the abyss". Is he running a lifestyle type company?

This reads like compensation theatrics.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#425
post #244

Earlier quoted context omitted.

You would do us all a service by telling us how your experiences of that have been.

I would say about 35% of the time I run into problems and eventually give up and go to GPT 5.5 and it much more efficiently handles the original task. Then I see the token costs going up and it motivates me to continue trying the open source ones.

There's going to be a tipping point where it's worth purchasing more hardware to run the next biggest size of the open model, if they show stepwise improvements that way.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#426
post #116

Earlier quoted context omitted.

Throwing out another factor: Chinese companies have been banned and/or limited from buying nvidia, and turned to local companies for their hardware. I haven't actually seen pricing/benchmarks comparing Chinese AI accelerators, but it wouldn't surprise me if that also worked out in their favor as well.

And, possibly, state subsidies at every level.

I have to point out the massive state subsidies in the united states for the tech companies and datacenter builders.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#427

Earlier quoted context omitted.

I'm using Deepseek-v4-pro as my main model and this is sometimes pretty annoying, I have to do some easy boring task, think "I'll just leave the agent to do it and go take a nap", but it's already done writing the code before I even walk away from the computer

DeepSeek is the fastest model in the benchmarks I've been doing ( https://swelljoe.com/post/will-it-mythos/ ). Followed not so closely by Opus 4.8 and even less closely by Gemini 3.5 Flash and GPT 5.5. I've been really impressed with it, so far. It's also among the best at doing the work, though still trailing the frontier models from Anthropic and OpenAI.

Nice benchmark, thanks! Which quants did you choose for the self hosted models?

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#428
post #103

Earlier quoted context omitted.

I remember when I had to wait minutes to get a high resolution image over a dialup connection. When computer and communications hardware advanced enough that I could get 30 high resolution images every second, there were brand new uses. In the case of LLMs, I could imagine that much faster operations allow you to introduce them as parts of systems that need to react to the real world at high speed, like factory equip…

The example in the video was a generation of a dashboard app of some sort. I can do that with a "normal speed" Claude in a few minutes. The difference is a few minutes. This is compared to a few weeks in old school development time. I don't have a problem with taking it a little "slow" (as in - few minutes) and lending my thought to it rather than just going for fast generation and who knows what's inside. I get your…

I frequently tell agent to do something, wait ~10 min (which is just enough that I can't/don't want to start anything else), ask it to change something, wait a few minutes again, and so on. So I'm basically idle while waiting for agent, and it would be great if it was faster.

It's like your compile times were ~10 min. Sure, it's not a huge deal, but it's sooo anoying

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#429

Earlier quoted context omitted.

They're leaving us in the dust on solar, while our current administration is still trying to put people in the ground to dig up more coal and die of black lung. https://en.wikipedia.org/wiki/Solar_power_in_China

They're building more coal than anyone. Also more nuclear than anyone, which one must assume you hate, because preferring solar requires you don't actually understand thing

Energy from coal in China decreased last year. The change is happening very quickly.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#430

Earlier quoted context omitted.

No, but nor can you keep track of what 10 agents are doing simultaneously. Hence the multitasking regret.

An agent can, you don't need to watch tasks, you can have a live digest with another tool.

Who watches the watchers?
Post reply on HN