Given that MiMo is as cheap as Deepseek ( previous discussion: https://news.ycombinator.com/item?id=48282814 ) multiplying that by 3x for ultra speed is still shockingly cheap.
MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
61–70 of 512 posts
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#62These price and speed optimization from Chinese providers, combined with the raising prices from American ones will change the game sooner than later. Many companies are finding issues with the AI bills already.
I see bigger problem with model inconsistency. You never know whether Anthropic will route your request to a cheaper model for the price of Opus. So you can never estimate how much a task will cost, because you might have to restart several times and pay for each attempt. Then you have to prompt models to gauge whether they are real or impostors which also adds to token usage.
For non subsidized plans? Pretty sure they'd need to put this in ToS, or law suites would have followed by now.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#63These price and speed optimization from Chinese providers, combined with the raising prices from American ones will change the game sooner than later. Many companies are finding issues with the AI bills already.
i've a Github copilot yearly subscription. Microsoft recently changed their billing to based on token. i'm still getting billed per premium request but GPT 5.4 is now 6x compare to 1x before.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#64Are you kidding me. Come back when you are ready for the users. I was hopping to try it, what a frustration.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#65Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#66Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#67This will be really powerful for voice. Being able to reason makes LLM so much smarter but with voice your latency budget is so tight that you can't spare the time typically.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#68I may sound like a shill, but exponential growth and all. We are going to get near instant software from prompt, multiple ones and then choose the best one. Discussions about choosing a library with the best syntactic sugar method naming is just as crazy as suggesting we type in assembly.
Sounds like exponential growth of crappy software. I'm not saying that before we didn't have mass produced crap in SE, but now it will turn into explosive overflow.
This strategy will seem to work really well until the economy that enabled that foundation to form is hollowed out. Then, there will be a reckoning (but we will have no choice but to march forth from there).
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#69I don't understand, given all they say, why this would not be made available to everyone at once? Why the limited release? They should have no trouble scaling it if it runs on a single rack.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#70I may sound like a shill, but exponential growth and all. We are going to get near instant software from prompt, multiple ones and then choose the best one. Discussions about choosing a library with the best syntactic sugar method naming is just as crazy as suggesting we type in assembly.
Sounds like exponential growth of crappy software. I'm not saying that before we didn't have mass produced crap in SE, but now it will turn into explosive overflow.