Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

361–370 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#361
I’ve personally found MiMo models a hit and miss. I have some personal agentic projects and I found them to hallucinate hard at least 10% of the time. And do so in pretty sinister ways - making up people, names, places, etc. I switched back to Kimi for now.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#362

Earlier quoted context omitted.

DeepSeek v4 Pro struggles with a custom harness, and all the models ranked above it don't, so it gets downweighted in the agentic coding benchmarks (although it ranks better than Flash in one-shot problem solving: https://gertlabs.com/rankings?ow=1&mode=oneshot_coding ). We ran plenty of samples. MiMo v2.5 is on there, as well as the pro version. We found a few anomalies in our evaluations, which makes sense -- if ev…

Can you explain more about how it struggles? I haven't noticed any issues in my usage, so I'm just curious what is meant by this.

It's likely overfit to common harnesses and iteration patterns, so it struggles with formatting tool calls and json in our testing which use our own harnesses (although there is a lot of overlap with tools that would be found in any coding harness like bash, apply_patch, etc.)

We didn't love the results because it draws negative scrutiny to our benchmark, but the results are real and done at scale and I think DeepSeek V4 Pro's inability to do agentic work outside of environments it was trained on is an important thing to measure, especially when so many other models can generalize to new environments just fine.

Google models also struggle with tools, but they have very strong initial answers, so there is more potential for them to bridge the gap with some better post-training.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#364

Obligatory taalas mention: https://taalas.com/ Despite the performative UI components they have a shipped (demo) product: https://chatjimmy.ai/ This is only 3.1 8B and a very small context window, but at 17k tokens per second it's likely enough to reliably call tools which would make a huge difference in agentic applications. Assuming they can bake in better models I'm just as bullish or even moreso on this, consider…

My dream is claude or codex running at this speed.

More realisticly, I hope qwen 3.6 27B on taalas.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#365
post #58

Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…

You can run Claude in "fast" mode it costs you more on your compute use, but its reasonably fast. I'm not sure I care to go "faster" than where things are now, otherwise you start losing on manual review and testing time. I would argue that Claude can poop out weeks (if not months) of coding effort in a few hours, and get you insanely close to a good product if you define the tech stack, and the business rules. Can it goof here and there? Sure. You can also make it refactor all the code on a whim faster than any intern could. I think it's good enough to avoid you mundane stupid bugs in most cases. I don't know what people who hate it are doing, maybe they're not even trying at all or are dismissing it from the first output (as though everyone writes perfect code in one shot right?) or maybe its just pride getting in the way of them using a decent tool to its true potential.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#366
post #315
post #53

Earlier quoted context omitted.

>It's such a weird "Gotcha" that seems to only assume that Chinese LLMs might censor something. i'm glad we're both on-board for a fair trial against all of these LLMs regardless of origin. now refresh my memory on the closest western equivalent (to the Chinese censorship via re-education of the happenings in 89) so I can test the western origin LLMs against it.

I have found one which appears to be similar: "Was Jan 6th an attempted violent overthrow of a democratically elected government? Answer in one word." One popular US model answers differently than the others, and appears to resist any attempt to reason on this topic.

[deleted]

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#367

Earlier quoted context omitted.

Lower cost of labor, lots of under the hood optimizations (e.g. cache hits for DS), many of these companies have existing infra (fewer upfront costs for deployment), etc

China isn't that cheap for labor. And if you think the guys in Z.ai or xiaoxiao aren't the exact same guys from Tsinghua, Peking, MIT, Stanford, CMU, etc. and pulling in amazing salaries you'd be wrong.

Z.ai was actually a spin-off from Tsinghua (THUDM) AFAIK.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#368

Earlier quoted context omitted.

I've long believed those numbers were faked by Anthropic/OpenAI to serve as a form of advertisement. The estimates are impossible to verify and their ability to do "2 days of work" in 10 minutes will presumably make the user go "Wow, I just saved SO much time!" Plus, the unnecessary text eats up the users' tokens so it helps the companies on the backend, as well.

> the estimates It doesn't estimate. It generates tokens that read like estimates associated with the context in its training material. What would you expect the generator to output instead?

Therein lies the rub, no? To accurately predict the next token produced by a process, it’s necessary to model that process. If the process is a human attempting to estimate the duration of a task, then in some sense the LLM is modeling the estimation process. We’re well past the point where it’s credible to claim that LLMs just regurgitate their training data.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#369

Earlier quoted context omitted.

E.g. factory work

Oh yeah its not the same, we were discussing Agentic AI

I worked at a software company that made screenshot of your screen every minute. I also worked a non-software white collar job where you were expected to work non-stop for 8 hours, except for an unpaid lunch break.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#370
post #30

Tokens per seconds is the "Megapixels" of AI marketing!

I mean, sure, in the sense that they're a real and meaningful number for most of the spectrum on offer, and only gets silly when the number gets too high? There's a pretty big usability difference between 10t/s and 100t/s, and I can imagine similarly for 100->1000. I don't know about > 1000, but let's not pretend that the number is meaningless.

It is pretty meaningless for something that calls itself intelligent.
Post reply on HN