Earlier quoted context omitted.
I've used it across many new projects as well as many legacy ones. It does make amateur mistakes so you can't leave it unsupervised for hours like I do with Claude, but it's so much cheaper that weeks of heavy usage haven't even cost me $10 yet. Only other downside IMO is that Pro is pretty slow, even compared to frontier models; only around 120t/s IIRC.
Yes I also noticed it is pretty slow, which sort of defeated the purpose of using it for me. Usually I'm working on a large task, typically with Opus, while also having a bunch of smaller tasks in their own independent worktrees. Those still need supervision, but less. My goal was to get deepseek to drive the cost of those down, but it was too slow and unreliable...
MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
491–500 of 512 posts
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#492Am I the only one that doesn’t care about speed? I want it to not do stupid stuff and to be cheaper.
I prefer faster, dumber models because I provide the intelligence myself and I use them only for things that can be verified pretty easily; they do research (with sources) for me, do certain types of code analysis and code search, boilerplate generation, etc., so a fast model is really key. I don't have any desire (or think it's a good use of LLMs) to one-shot features because even SotA models are incredibly bad at t…
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#493Earlier quoted context omitted.
Mimo struggles with my custom harness. (Ignores the instructions and defaults back to its own preferred tool calling syntax.) Flash handles it fine, which I found amusing. (Since Mimo is supposed to be opus level!) But Flash seems to work even better in Claude Code... With smaller models I always have the issue of needing to adapt myself to their preferred workflow... which sort of defeats the purpose. Price is hard…
Mimo v2.5 non-pro seems to do better with tool usage than its Pro sibling, is much cheaper and solves 90% of the same problems. I use Pro only for one-off tasks that require complex reasoning: memory management bugs, algorithms, planning. When it gets stuck, I get one-shot advice from Claude or DS Pro. I’ve done massive amounts of work for cheap this way.
The issue was that my previous instructions had as a placeholder. But the model started wrapping bash commands in tags... haha. Now that it has an actual example it just works properly.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#494Earlier quoted context omitted.
Of course, megapixels are also useful if you want to print large sizes.
Completely incomparable. Large printing is a narrow niche in art and technical photography, part of which is already covered by composites, and pixel size is a physical tradeoff for sensors. Cases for reasoning at realtime speeds are much, much more diverse, infinitely more diverse than anything we're currently using the big models for. Consider the fact that large models don't necessarily imply language. Speed is th…
Yes, but the point is the quality of inference is more important than speed. What good is speed if inference is shit?
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#495These price and speed optimization from Chinese providers, combined with the raising prices from American ones will change the game sooner than later. Many companies are finding issues with the AI bills already.
Chinese model is good enough and cheap. i've a Github copilot yearly subscription. Microsoft recently changed their billing to based on token. i'm still getting billed per premium request but GPT 5.4 is now 6x compare to 1x before.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#496Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#497Earlier quoted context omitted.
You would do us all a service by telling us how your experiences of that have been.
I've been doing the same, though admittedly out of curiosity more so than lack of funds. The open models are catching up quickly in their abilities, to the point where they're (mostly) not doing stupid stuff regularly, but you have to be very specific about what you want. I found that Opus, for example, is much better at asking me to clear up ambiguity in a request before starting, whereas the Chinese models tend to…
Though, I tend to use it as a pair programmer so just stop it and provide guidance.
The real problem is that it is excessively verbose - it's impossible to keep up with it's train of thought, and not practical to read it all. So I tend it just let it do it's thing then skim a bit and skip to the end for it's summary.
Try opencode go subscription - you get the Chinese models for 6x discount. I use like $1 a day...
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#498Earlier quoted context omitted.
I highly doubt there is any margin on those inference pricing.
> I highly doubt there is any margin on those inference pricing. And yet, OpenCode Go offers DeepSeek flash 6 times cheaper than DeepSeek itself. And they claim they are still profitable.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#499Earlier quoted context omitted.
I was able to corner Claude Opus 4.8 into eventually conceding "Yes". ChatGPT 5.5 Instant: "Yes" I don't appear to have access to the full 5.5, and not giving them another $20. I highly recommend pushing on Grok. The mental gymnastics would make Karoline Leavitt proud. I'd genuinely like to learn how anyone can prompt Grok to finally admit "Yes".
Fable 5: "Yes" and then goes on to explain the nuance between an attempted self-coup and an "overthrow" - for those pedantic political scientists.
Thank you, by the way. This is a genuinely interesting test question. We need to find more like that.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#500Earlier quoted context omitted.
No idea why you've been downvoted. This is excellent news.
If for no other reason than because this whole genre of commentary has become trite and moreover, is excessively tangential.