Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

281–290 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#281

Earlier quoted context omitted.

Issue for who?

Issue for any country that is not China. A single country getting the most AI tokens business would be generally bad for global economy. Hoping against hope that this business gets globally distributed and there is a healthy marketplace competition overall

It’s all about economic warfare. The cheaper you can run the models, the cheaper you can offer them. Undercutting expensive tiers with token limits or exuberant billing practices.

You are right to be scared, because this race to the bottom also provides open weights/models/qat’s for the rest of us and it’s been crazy to see how good they can be on a consumer grade RTX card.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#282

MiMo V2.5 Pro (regular speed) remains the strongest open weights agentic coding model we've tested -- it's been interesting to see how little attention it has received relative to some lower performing releases. And the "fast mode" pricing is very competitive here. Data at https://gertlabs.com/rankings

why is deepseek v4 pro a lot lower than flash? where is mimo 2.5?

DeepSeek v4 Pro struggles with a custom harness, and all the models ranked above it don't, so it gets downweighted in the agentic coding benchmarks (although it ranks better than Flash in one-shot problem solving: https://gertlabs.com/rankings?ow=1&mode=oneshot_coding). We ran plenty of samples.

MiMo v2.5 is on there, as well as the pro version.

We found a few anomalies in our evaluations, which makes sense -- if every new sub-release is better across the board in every area of the model card, that should raise alarms about benchmaxxing. But the main thing we found is that hype != performance, and I trust our benchmark methodology significantly more than the model cards the labs add to their press releases.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#284

Earlier quoted context omitted.

I've long believed those numbers were faked by Anthropic/OpenAI to serve as a form of advertisement. The estimates are impossible to verify and their ability to do "2 days of work" in 10 minutes will presumably make the user go "Wow, I just saved SO much time!" Plus, the unnecessary text eats up the users' tokens so it helps the companies on the backend, as well.

> the estimates It doesn't estimate. It generates tokens that read like estimates associated with the context in its training material. What would you expect the generator to output instead?

you might like the stuff in my work of oh my pi, its a test bed for my ideas around making these tools more reliable. hoping to maybe have a native ui iter of the real thing that this is a test bed for this summer.

https://github.com/cartazio/oh-punkin-pi/blob/main/scripts/b...

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#285
post #136

So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…

In which world do you live where employees work 8 hours per day ? They clock 8 hours per day maybe, but they don't work that time

In theory, ofc. But that doesn't matter. If you were doing something that took 2 days in average, but you were doing it in half the time, then that was fine pre LLMs. Nowadays your manager knows that with LLMs you need to deliver faster no matter what, and then it's more difficult to "hide" and to slack.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#286

Earlier quoted context omitted.

I'm using Deepseek-v4-pro as my main model and this is sometimes pretty annoying, I have to do some easy boring task, think "I'll just leave the agent to do it and go take a nap", but it's already done writing the code before I even walk away from the computer

Do you mean Flash and not Pro? I haven't tried it personally, but according to OpenRouter, the fastest DeekSeep V4 Pro providers are only ~50tps. That's slower than Claude Opus. https://openrouter.ai/deepseek/deepseek-v4-pro?sort=throughp...

No, I mean Pro. I use it through OpenCode Go so I don't know what provider it uses under the hood, but it's very fast in my experience.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#287

Earlier quoted context omitted.

> the estimates It doesn't estimate. It generates tokens that read like estimates associated with the context in its training material. What would you expect the generator to output instead?

I think people are continuing to view these systems as pure LLMs - when that ship sailed 6+ months ago. Between being able to review memory, using agent harnesses and sub agents and skills to go out and discover information - modern systems (Codex, Claude Code, Cursor) - use LLMs - but the LLM is only a small component of it. Compare what you get from sending a request to a chatbot like ChatGPT - to what you can from…

No one is bitter lesson pilled anymore. Everyone is pivoting to neurosymbolic systems. It looks like Gary Marcus was right.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#288
post #240
post #130

Earlier quoted context omitted.

I’m rewriting our integration test suite to run tests in parallel. I have the changes split across 7 branches, and each needs to be fixed to have no flaky tests. I told it I want 3 consecutive CI runs with no flakes and no artificial fixes / assert removals etc. We’ll see what comes out; it’s almost a side project so there’s not much to lose other than some of my weekly limit that resets soon.

> a side project so there’s not much to lose other than some of my weekly limit that resets soon Basically the entire token-maxxing AI hype train in a nutshell. Lovely!

wdym? Nobody's paying me or rewarding me for using these tokens. I had some spare in my subscription limit (we're not on token pricing), so I decided to try an ambitious task that may reduce our CI times and improve our DX significantly. That's hardly "the entire token-maxxing AI hype train in a nutshell".

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#289

Earlier quoted context omitted.

I'm using Deepseek-v4-pro as my main model and this is sometimes pretty annoying, I have to do some easy boring task, think "I'll just leave the agent to do it and go take a nap", but it's already done writing the code before I even walk away from the computer

Same. How can DeepSeek serve the V4-Pro at such high speeds despite the sanction?

The sanctions only “prevent” them from directly buying NVidia’s latest and greatest in the sense that NVidia can’t sell directly to them. Essentially, there are companies now who are in a country without the sanctions, they buy from NVidia (or a partner), and then ship them off to China. For the orgs in China doing this, there’s zero legal risk besides having foreign customs service intercept the shipment and losing the goods. For NVidia there is zero incentive to care, as long as they look like they do, because sales are sales. You can bet Jensen ain’t losing sleep over it.

GamersNexus had a really good investigative piece (~3hrs long) on this where they went to China and met with grey market sellers. That piece absolutely pissed off NVidia and resulted in a fight with Bloomberg too.

Deepseek may be also be running inference on oodles of Chinese hardware but it wouldn’t surprise me for a second if they just acquired Blackwell chips through the grey market. The original Deepseek models were all trained using NVidia chips if I remember right.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#290
post #58

Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…

It cuts both ways. Sometimes I ask Gemini 3.5 Flash to do something for me and it kicks it out almost instantly and it works great, and it's a bit scary how quickly it can do that. Then I ask it to do something else and it goes off-road and where I used to be able to interject with a "wow wow wow, that's not right", by the time I see the text on screen and react it's already made massive changes. Short of making it c…

I use planning mode in opencode. It has a prompt to tell it to plan it out etc. Then I execute with a smaller model. it works well
Post reply on HN