Live data from Hacker News

Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%

platform.xiaomimimo.com

101–110 of 165 posts

Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%

#101
post #88

Earlier quoted context omitted.

> US AI companies have no chance of recouping even fraction of their valuations. A big caveat here is that many US companies (particularly in sensitive industries, like defense) will likely not want to (or not be allowed to) use Chinese models for anything of substance.

What about self host Chinese models?

Still no.

Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%

#102
post #46
post #29

That's deliberate. US AI companies have no chance of recouping even fraction of their valuations. PS: Have not tried this but Deepseev4 Flash (not even Deepseekv4 Pro version) with set to "high" has pretty much Claud Opus 4.7 level of capabilities and is lightening fast and dirty cheap. Hours and hours of conversation barely costs few cents.

I am very happy with DSv4 for their price/performance but neither of them are comparable to Opus.

Yeah, I really like and use DSv4Pro for personal projects, but I also use Opus all the time at work and they are definitively not at the same level.

I can only conclude that people who claim they are aren't doing anything close to the edge of what these models are capable of or any niche things.

I would say DSv4Pro is around the same level as Sonnet.

Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%

#104

Insane. 3 points behind opus on the artificialanalysis index. Mimo cost ~$400 at the old price, so about $40 today. Opus cost ~$5000 That's over 100x cheaper, and just 3 points behind. I can't wait to experiment with an llm consortium of 100 deepseek and mimo models. Crazy times. Shut up and take my m̶o̶n̶e̶y̶ data! Edit: Gemini on google search told me I could write strikethrough text on hn using . Mimo told me it w…

benchmarks we deserve: google search quick ai answers vs full llm model :)

Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%

#105
post #98
post #91

Hot take: The reason this is happening is because the market for Chinese AI models hosted by Chinese companies is struggling. Even the market for Chinese AI models hosted by western companies is soft: During the week of May 18, OpenRouter processed 3.4T DeepSeek v3 Flash tokens (their most popular model). Google has announced that Gemini is processing 746T per week; Claude is probably processing more. And the Chinese…

> OpenRouter processed 3.4T DeepSeek v3 Flash > Gemini is processing 746T per week I read this totally differently. A startup nobody really knows is doing half a percent of Google on a commodity task?!? Google, which puts Gemini on billions of devices by default, without the user asking? Google, which is distributing Gemini to users who are unaware they are even using it? Versus a startup that does not even have a lo…

Unfortunately, the market doesn't generally let you buy Blackwells with "we got half a percent of Google's marketshare with a model we're literally giving away for free [1]". You need that thing we call Capital. But, they may certainly opt to have it written on their gravestone, as Google is (checks notes) continuing to put Gemini on billions of devices and doing quadrillions of tokens per month.

[1] https://openrouter.ai/deepseek/deepseek-v4-flash:free

Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%

#106
post #2

OK. Google was just killed. How is it possible to reduce the price by 99%??????? This is crazy

The reduction is in cached inputs. I've commented about this before but many labs, except Deepseek and Xaomi now, absolutely scam you for cached reads. You are basically paying out the nose for a few seconds of VRAM residence if you are giving significant money for cache reads. The very nature of autoregressive language modeling is that every single output token produced "reads" the cache. So in principle the price f…

No one is producing one output token though.

And using up gpus for that cache is a pretty big opportunity cost. I highly doubt it's done in vram. That would be insane for the one hour caches.

So its memory + the time it takes to unload/load into vram + the extra cost per output token

Is it a scam? Idk

Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%

#107
post #99

Earlier quoted context omitted.

> there's obviously a question on how/why are they lowering the prices so much. Same reason they release some of the models for free: They are trying to capture market share.

LLM providers can't "capture" anything. People loved Claude Code because it was cheap and good. Not cheap anymore? People switching to Codex, DS4 etc. Their only moat is maybe being SOTA but that only lasts so long before everyone else catches up.

I mean there is a minor moat. Most people don't enjoy switching providers or models. If you can get people to trust you'll stay near frontier, they'll stick around even when you aren't the best. Claude is a prime example of this

Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%

#109
post #91

Hot take: The reason this is happening is because the market for Chinese AI models hosted by Chinese companies is struggling. Even the market for Chinese AI models hosted by western companies is soft: During the week of May 18, OpenRouter processed 3.4T DeepSeek v3 Flash tokens (their most popular model). Google has announced that Gemini is processing 746T per week; Claude is probably processing more. And the Chinese…

OpenRouter is not indicative of volume. Most high volume clients will go to the providers directly. There's not point to paying the 5% OR cut if you know what you want.

Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%

#110
post #105
post #98

Earlier quoted context omitted.

> OpenRouter processed 3.4T DeepSeek v3 Flash > Gemini is processing 746T per week I read this totally differently. A startup nobody really knows is doing half a percent of Google on a commodity task?!? Google, which puts Gemini on billions of devices by default, without the user asking? Google, which is distributing Gemini to users who are unaware they are even using it? Versus a startup that does not even have a lo…

Unfortunately, the market doesn't generally let you buy Blackwells with "we got half a percent of Google's marketshare with a model we're literally giving away for free [1]". You need that thing we call Capital. But, they may certainly opt to have it written on their gravestone, as Google is (checks notes) continuing to put Gemini on billions of devices and doing quadrillions of tokens per month. [1] https://openrout…

This is a bizarre comment for a couple of reasons.

First, obviously everyone involved understands that someone has to pay to provide a free service. Everyone involved also knows that this sometimes makes sense as a business strategy (I have not paid to ship anything from Amazon for close to two decades).

Second, OpenRouter's business model specifically does not require them to run all (any?) of the models available through the platform. Provider is one of the choices when you choose a model, and each provider can have separate pricing.

The link you posted shows only one provider, Crucible. That may/may not be affiliated with OpenRouter? Even assuming an affiliation, it's opaque who is subsidizing this usage. Is it OpenRouter or Crucible?

All of this is somewhat of a distraction. Even if someone gave search away for free (like Google), it would still be an accomplishment to get to half a percent of Google's volume. Or to sell half a percent of the volume of Android phones. Or whatever.

Kudos to the OpenRouter team!

Post reply on HN