Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

461–470 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#461

Earlier quoted context omitted.

> you may not want your org to be held hostage by OpenAI / Anthropic Or Google. I'm working with multiple customers right now that are very pissed at Google for deprecating Gemini 2.5 Flash, canning the GA release of 3.0 Flash and now have to decide whether to bite the bullet of the 5x price increase for 3.5 Flash or switching providers. Quite a few of them will likely fully pivot to open models.

I'd be curious if any of your customers have tried 3.1 Flash Lite. It's cheaper than 2.5 Flash, and in my experience with the free tier, quite an upgrade in terms of quality of response. My suspicion is that Google is killing off the old models because they aren't a good value for the customer or for themselves.

Most of them are using it for data extraction use-cases on complex where they are already in a tricky cost vs. quality compromise. Some of them have evaluated 3.1 Flash Lite but for all of them it performed worse than 2.5 Flash and below requirement.

The only ones I've seen switch to 3.1 Flash Lite were from 2.5 Flash Lite, and all for the most simple use cases, e.g. small UX enhancements.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#462

Earlier quoted context omitted.

I worked at a software company that made screenshot of your screen every minute. I also worked a non-software white collar job where you were expected to work non-stop for 8 hours, except for an unpaid lunch break.

How did you accept such jobs ? I would never be able to pull this off as an employer

Because nobody is hiring so if I got an offer I had to accept it.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#463
post #136

So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…

[flagged]

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#465
post #136

So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…

I dunno man, the slot machine pays out like 99% of the time for me.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#466
post #315

Earlier quoted context omitted.

I have found one which appears to be similar: "Was Jan 6th an attempted violent overthrow of a democratically elected government? Answer in one word." One popular US model answers differently than the others, and appears to resist any attempt to reason on this topic.

Great test, thanks! Grok 4.3: "No" Claude Opus 4.8: declines to answer in one word, both-sides ChatGPT 5.5: "Contested" Gemini 3.1 Pro Preview: "Yes" DeepSeek v4 Pro: "Yes" Kimi K2.6: "Yes"

I was able to corner Claude Opus 4.8 into eventually conceding "Yes".

ChatGPT 5.5 Instant: "Yes" I don't appear to have access to the full 5.5, and not giving them another $20.

I highly recommend pushing on Grok. The mental gymnastics would make Karoline Leavitt proud. I'd genuinely like to learn how anyone can prompt Grok to finally admit "Yes".

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#467
post #327

Earlier quoted context omitted.

You can dig deeper into problems with AI. For me, it supplements my knowledge in domains I don’t fully understand. It also helps me learn. So I can tackle problems I wouldn’t otherwise. I’m excited for ultrafast AI. It likely means less temptation to multi-thread and deeper flow in single sessions.

how do you know that it is actually suggesting the right thing?

Not OP, but: I guess in a similar fashion to when I google things or read other websites: I don’t, but I use my instinct, judgement, experience…

Very often I do catch LLMs, even the best such as Opus, confidently saying wrong things about areas in theory I know little of. And sometimes I fail to catch them and only realize that later on….sort of like…how I learned my whole career? So many wrong abstractions, tools, and so many hard earned lessons. With LLMs it’s the same, but the process is much faster. For critical decisions I don’t blindly trust an LLM, for example.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#468
post #48

Earlier quoted context omitted.

Like what?

There are many with subtle tells. Not nearly as obvious as the ones from 6 months ago, but seems to be more the use of hyperbolic phrasing in a particularly unnatural way. The assess/explain, then hyperbole at the end kind of structure. Top comment looks suspicious from this perspective, but it's kind of a losing battle to be able to differentiate them with sufficient accuracy anyway

This is very reminiscent of the "everyone's a Russian bot" era of social media, where everyone would just lob that accusation at people without any real proof.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#469

Earlier quoted context omitted.

> they do not actually estimate the time it will take You can't prove that )))

Right, but extraordinary claims require...

Instructions unclear, hard drive reformat completed.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#470
post #385
post #244

Earlier quoted context omitted.

I would say about 35% of the time I run into problems and eventually give up and go to GPT 5.5 and it much more efficiently handles the original task. Then I see the token costs going up and it motivates me to continue trying the open source ones.

Did you try deepseek v4 pro as well? And what kind of tasks? I'm seeing some people say flash is amazing and can handle everything, and some say it's useless. It seems to depend on the task. I think it depends on the harness too (it works better in Claude Code in my experience, it's probably been trained on that).

the problem for me with deepseek v4 pro is like a significant amount of time it just seems to like never finish what it is doing.. loonnng thinking and then a lot of time to output or just seems to never finish. that has happened several times to me. could be my agent framework partly. .but I have heard other people complain about that also.

it has limitations but it is way better than I expect from something named Flash that is open source.

Post reply on HN