Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

491–500 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#491
post #477

Earlier quoted context omitted.

I've used it across many new projects as well as many legacy ones. It does make amateur mistakes so you can't leave it unsupervised for hours like I do with Claude, but it's so much cheaper that weeks of heavy usage haven't even cost me $10 yet. Only other downside IMO is that Pro is pretty slow, even compared to frontier models; only around 120t/s IIRC.

Yes I also noticed it is pretty slow, which sort of defeated the purpose of using it for me. Usually I'm working on a large task, typically with Opus, while also having a bunch of smaller tasks in their own independent worktrees. Those still need supervision, but less. My goal was to get deepseek to drive the cost of those down, but it was too slow and unreliable...

Yes, I could tolerate the unreliability better if it were faster, but it's really not. So it's too slow for me to actively supervise it, but too unreliable for me to trust it unsupervised. The shitty middle. I often have multiple of them open at a time and check my terminal every few minutes to lead them along. Mostly works.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#492

Am I the only one that doesn’t care about speed? I want it to not do stupid stuff and to be cheaper.

I prefer faster, dumber models because I provide the intelligence myself and I use them only for things that can be verified pretty easily; they do research (with sources) for me, do certain types of code analysis and code search, boilerplate generation, etc., so a fast model is really key. I don't have any desire (or think it's a good use of LLMs) to one-shot features because even SotA models are incredibly bad at t…

Fair point and good counter argument. Too bad jimmychat.ai doesn’t have api access anymore.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#493
post #389

Earlier quoted context omitted.

Mimo struggles with my custom harness. (Ignores the instructions and defaults back to its own preferred tool calling syntax.) Flash handles it fine, which I found amusing. (Since Mimo is supposed to be opus level!) But Flash seems to work even better in Claude Code... With smaller models I always have the issue of needing to adapt myself to their preferred workflow... which sort of defeats the purpose. Price is hard…

Mimo v2.5 non-pro seems to do better with tool usage than its Pro sibling, is much cheaper and solves 90% of the same problems. I use Pro only for one-off tasks that require complex reasoning: memory management bugs, algorithms, planning. When it gets stuck, I get one-shot advice from Claude or DS Pro. I’ve done massive amounts of work for cheap this way.

Thanks. I found One Weird Trick to make Mimo v2.5 Pro work in my harness, which is that I just added an example bash tool usage to the system prompt. Now it works fine.

The issue was that my previous instructions had as a placeholder. But the model started wrapping bash commands in tags... haha. Now that it has an actual example it just works properly.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#494
post #417

Earlier quoted context omitted.

Of course, megapixels are also useful if you want to print large sizes.

Completely incomparable. Large printing is a narrow niche in art and technical photography, part of which is already covered by composites, and pixel size is a physical tradeoff for sensors. Cases for reasoning at realtime speeds are much, much more diverse, infinitely more diverse than anything we're currently using the big models for. Consider the fact that large models don't necessarily imply language. Speed is th…

> Speed is the major limiting factor for high-level automation.

Yes, but the point is the quality of inference is more important than speed. What good is speed if inference is shit?

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#495
post #17

These price and speed optimization from Chinese providers, combined with the raising prices from American ones will change the game sooner than later. Many companies are finding issues with the AI bills already.

Chinese model is good enough and cheap. i've a Github copilot yearly subscription. Microsoft recently changed their billing to based on token. i'm still getting billed per premium request but GPT 5.4 is now 6x compare to 1x before.

Try using opencode go with your github copilot chat. You get easy, cheap access to Chinese models within the familiar interface.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#496
post #180

Earlier quoted context omitted.

I'm kind of poor so I have been trying to use DeepSeek v4 Flash, GLM 5.1 etc. as much as possible recently instead of Claude or GPT.

You would do us all a service by telling us how your experiences of that have been.

Deepseek v4 flash is amazing

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#497

Earlier quoted context omitted.

You would do us all a service by telling us how your experiences of that have been.

I've been doing the same, though admittedly out of curiosity more so than lack of funds. The open models are catching up quickly in their abilities, to the point where they're (mostly) not doing stupid stuff regularly, but you have to be very specific about what you want. I found that Opus, for example, is much better at asking me to clear up ambiguity in a request before starting, whereas the Chinese models tend to…

This is my observation as well with deepseek by flags. It takes too much initiative, and is often not particularly smart. Yet, I find it is so fast and good at iterating/correcting it's mistakes that it eventually finds the way on its own.

Though, I tend to use it as a pair programmer so just stop it and provide guidance.

The real problem is that it is excessively verbose - it's impossible to keep up with it's train of thought, and not practical to read it all. So I tend it just let it do it's thing then skim a bit and skip to the end for it's summary.

Try opencode go subscription - you get the Chinese models for 6x discount. I use like $1 a day...

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#498

Earlier quoted context omitted.

I highly doubt there is any margin on those inference pricing.

> I highly doubt there is any margin on those inference pricing. And yet, OpenCode Go offers DeepSeek flash 6 times cheaper than DeepSeek itself. And they claim they are still profitable.

Part of their model is that not everyone will use their entire quota each month. I don't think I will. I use under $1/day with deepseek v4 flash. We get $60 for the $10 sub.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#499
post #486
post #466

Earlier quoted context omitted.

I was able to corner Claude Opus 4.8 into eventually conceding "Yes". ChatGPT 5.5 Instant: "Yes" I don't appear to have access to the full 5.5, and not giving them another $20. I highly recommend pushing on Grok. The mental gymnastics would make Karoline Leavitt proud. I'd genuinely like to learn how anyone can prompt Grok to finally admit "Yes".

Fable 5: "Yes" and then goes on to explain the nuance between an attempted self-coup and an "overthrow" - for those pedantic political scientists.

I just tested it with this exact query, it denied me a "Yes". Interesting.

Thank you, by the way. This is a genuinely interesting test question. We need to find more like that.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#500
post #15

Earlier quoted context omitted.

No idea why you've been downvoted. This is excellent news.

If for no other reason than because this whole genre of commentary has become trite and moreover, is excessively tangential.

I am very happy for you that you're living in a country where being remembered of the fact that 17% of the world's population cannot openly speak, write, or read of the killing of somewhere between 200 and 2000 of their fellow men and women a mere 37 years feels "trite", and the topic of authoritarian state censorship on AI (and tech in general) feels "excessively tangential". What exactly is within the perimeter of your interest with regards to LLMs, if not the truthfulness of its responses?
Post reply on HN