Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

191–200 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#191

MiMo V2.5 Pro (regular speed) remains the strongest open weights agentic coding model we've tested -- it's been interesting to see how little attention it has received relative to some lower performing releases. And the "fast mode" pricing is very competitive here. Data at https://gertlabs.com/rankings

why is deepseek v4 pro a lot lower than flash? where is mimo 2.5?

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#192

Earlier quoted context omitted.

I am more and more inclined into not believing this crappy software theory. Especially as teams invest in proper agentic harnessing. We have had a champion in our team that has invested a lot of time into it over the last 4 months, and if anything, quality has improved, not decreased. Architecture is more coherent, codebase has been cleaned up, agents find information quickly, code produced is very solid and my role…

It makes no sense. I mean, T2 covered this: "Watching John with the machine, it was suddenly so clear. The terminator would never stop. It would never leave him, and it would never hurt him, never shout at him, or get drunk and hit him, or say it was too busy to spend time with him. It would always be there. And it would die to protect him. Of all the would-be fathers who came and went over the years, this thing, thi…

I'm confused, what does not make sense?

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#194

Earlier quoted context omitted.

Because this never gets brought up about US models, which have just as much censorship as the Chinese ones.

US models are happily parroting Russian fakes. US censorship is a joke.

Can you point me to one example? (Without web search, of course). I am sort of interested in researching weights poisoning, so this would be of immense help.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#195

Earlier quoted context omitted.

Please educate us - which accurate and provable events in history are censored by US based LLMs as part of a government enforced reeducation campaign?

Does it even matter which agendas get censored? Like why won't my Claude tell me how to make sarin gas? I'd genuinely like to understand it. Sure, you can always reach for a justification saying "preventing terrorism" but the same argument can be made by Chinese AI labs. What actually matters is that the mere tool is withholding information at all, and that the boundaries were set by whoever designed it. Dont get me…

You can read this in Wikipedia. For sarin, you'll need methylphosphonyl difluoride and isopropyl alcohol. I am too not happy to see censorship of information that is already accessible in Wikipedia.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#196

Earlier quoted context omitted.

Agent mania setting in It's also pretty funny sometimes how it gives weird future roadmap estimates ("part 2 - 3 weeks, part 3 - 2 months", etc.) and when you tell it to actually do those changes it's pretty much done in half an hour

I've long believed those numbers were faked by Anthropic/OpenAI to serve as a form of advertisement. The estimates are impossible to verify and their ability to do "2 days of work" in 10 minutes will presumably make the user go "Wow, I just saved SO much time!" Plus, the unnecessary text eats up the users' tokens so it helps the companies on the backend, as well.

I agree with you that labs are benefiting from those outputs but I'm skeptical that labs are purposefully training the models to produce those outputs.

Raw pre-training data includes plenty of conversations between professional builders and some of those include estimates.

I believe the outputs are a training coincidence with consequences that are opportunitistic for the labs.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#197

Earlier quoted context omitted.

It's going to be an issue when China ends up scaling faster as well. Faster tokens, faster clusters, qat models, fp4, it's getting scary.

Issue for who?

American Politics and the far right.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#198

I don't understand, given all they say, why this would not be made available to everyone at once? Why the limited release? They should have no trouble scaling it if it runs on a single rack.

It uses significantly more resources obviously. And/or they have to configure or reconfigure servers for it, which takes time, and doesn't make sense until they have proven the demand at the higher price point.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#199
post #53
post #23

Earlier quoted context omitted.

Can I ask an honest question? Why does that matter in the slightest? LLMs come out with completely incorrect information all the time, and Western LLMs are censored for various topics too. It's such a weird "Gotcha" that seems to only assume that Chinese LLMs might censor something.

>It's such a weird "Gotcha" that seems to only assume that Chinese LLMs might censor something. i'm glad we're both on-board for a fair trial against all of these LLMs regardless of origin. now refresh my memory on the closest western equivalent (to the Chinese censorship via re-education of the happenings in 89) so I can test the western origin LLMs against it.

the civil war was only ever and exclusively about states rights

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#200
post #136

So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…

I dig into problems way, way deeper with AI than without. I can also add a lot more polish to features, add more test coverage, write more documentation, explore multiple approaches rather than go with gut-feel, and so on.
Post reply on HN