MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
361–370 of 512 posts
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#362Earlier quoted context omitted.
DeepSeek v4 Pro struggles with a custom harness, and all the models ranked above it don't, so it gets downweighted in the agentic coding benchmarks (although it ranks better than Flash in one-shot problem solving: https://gertlabs.com/rankings?ow=1&mode=oneshot_coding ). We ran plenty of samples. MiMo v2.5 is on there, as well as the pro version. We found a few anomalies in our evaluations, which makes sense -- if ev…
Can you explain more about how it struggles? I haven't noticed any issues in my usage, so I'm just curious what is meant by this.
We didn't love the results because it draws negative scrutiny to our benchmark, but the results are real and done at scale and I think DeepSeek V4 Pro's inability to do agentic work outside of environments it was trained on is an important thing to measure, especially when so many other models can generalize to new environments just fine.
Google models also struggle with tools, but they have very strong initial answers, so there is more potential for them to bridge the gap with some better post-training.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#363This could bring proper desktop AI to the average laptop user, which could be a game changer for running local models.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#364Obligatory taalas mention: https://taalas.com/ Despite the performative UI components they have a shipped (demo) product: https://chatjimmy.ai/ This is only 3.1 8B and a very small context window, but at 17k tokens per second it's likely enough to reliably call tools which would make a huge difference in agentic applications. Assuming they can bake in better models I'm just as bullish or even moreso on this, consider…
My dream is claude or codex running at this speed.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#365Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#366Earlier quoted context omitted.
>It's such a weird "Gotcha" that seems to only assume that Chinese LLMs might censor something. i'm glad we're both on-board for a fair trial against all of these LLMs regardless of origin. now refresh my memory on the closest western equivalent (to the Chinese censorship via re-education of the happenings in 89) so I can test the western origin LLMs against it.
I have found one which appears to be similar: "Was Jan 6th an attempted violent overthrow of a democratically elected government? Answer in one word." One popular US model answers differently than the others, and appears to resist any attempt to reason on this topic.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#367Earlier quoted context omitted.
Lower cost of labor, lots of under the hood optimizations (e.g. cache hits for DS), many of these companies have existing infra (fewer upfront costs for deployment), etc
China isn't that cheap for labor. And if you think the guys in Z.ai or xiaoxiao aren't the exact same guys from Tsinghua, Peking, MIT, Stanford, CMU, etc. and pulling in amazing salaries you'd be wrong.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#368Earlier quoted context omitted.
I've long believed those numbers were faked by Anthropic/OpenAI to serve as a form of advertisement. The estimates are impossible to verify and their ability to do "2 days of work" in 10 minutes will presumably make the user go "Wow, I just saved SO much time!" Plus, the unnecessary text eats up the users' tokens so it helps the companies on the backend, as well.
> the estimates It doesn't estimate. It generates tokens that read like estimates associated with the context in its training material. What would you expect the generator to output instead?
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#369Earlier quoted context omitted.
E.g. factory work
Oh yeah its not the same, we were discussing Agentic AI
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#370Tokens per seconds is the "Megapixels" of AI marketing!
I mean, sure, in the sense that they're a real and meaningful number for most of the spectrum on offer, and only gets silly when the number gets too high? There's a pretty big usability difference between 10t/s and 100t/s, and I can imagine similarly for 100->1000. I don't know about > 1000, but let's not pretend that the number is meaningless.