Earlier quoted context omitted.
> there's obviously a question on how/why are they lowering the prices so much. Same reason they release some of the models for free: They are trying to capture market share.
LLM providers can't "capture" anything. People loved Claude Code because it was cheap and good. Not cheap anymore? People switching to Codex, DS4 etc. Their only moat is maybe being SOTA but that only lasts so long before everyone else catches up.
Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
111–120 of 165 posts
Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
#112Anything to destroy US tech companies is welcome. They aren't aiming companies but users which many have no common sense and grant these agentic AI access to everything. All the restrictions the US imposed to CH, will be reverted back and it will be even worse, because now the data is not reaching the US gov ( we all know they have access to US big techs data ) but CH. I really hope this goes viral and breaks Nvidia/…
Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
#113Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
#114Hot take: The reason this is happening is because the market for Chinese AI models hosted by Chinese companies is struggling. Even the market for Chinese AI models hosted by western companies is soft: During the week of May 18, OpenRouter processed 3.4T DeepSeek v3 Flash tokens (their most popular model). Google has announced that Gemini is processing 746T per week; Claude is probably processing more. And the Chinese…
OpenRouter is not indicative of volume. Most high volume clients will go to the providers directly. There's not point to paying the 5% OR cut if you know what you want.
But you're right that OpenRouter is only one data point. It is, unfortunately, one of the few we have.
Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
#115Earlier quoted context omitted.
Unfortunately, the market doesn't generally let you buy Blackwells with "we got half a percent of Google's marketshare with a model we're literally giving away for free [1]". You need that thing we call Capital. But, they may certainly opt to have it written on their gravestone, as Google is (checks notes) continuing to put Gemini on billions of devices and doing quadrillions of tokens per month. [1] https://openrout…
This is a bizarre comment for a couple of reasons. First, obviously everyone involved understands that someone has to pay to provide a free service. Everyone involved also knows that this sometimes makes sense as a business strategy (I have not paid to ship anything from Amazon for close to two decades). Second, OpenRouter's business model specifically does not require them to run all (any?) of the models available t…
My original DeepSeek v4 Flash token counts spanned all providers of that model, both paid and free; I merely pointed out the free provider to substantiate a point that DeepSeek's product may be so bad that they could quite literally give it away and people would still prefer to pay (a lot) to OpenAI, Anthropic, or Google. Why this is the case, I leave as a exercise to the reader; I'm just citing numbers and facts.
Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
#116Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
#117Anything to destroy US tech companies is welcome. They aren't aiming companies but users which many have no common sense and grant these agentic AI access to everything. All the restrictions the US imposed to CH, will be reverted back and it will be even worse, because now the data is not reaching the US gov ( we all know they have access to US big techs data ) but CH. I really hope this goes viral and breaks Nvidia/…
fun fact, CH is an ISO code of Switzerland, and China is CN
Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
#118Hot take: The reason this is happening is because the market for Chinese AI models hosted by Chinese companies is struggling. Even the market for Chinese AI models hosted by western companies is soft: During the week of May 18, OpenRouter processed 3.4T DeepSeek v3 Flash tokens (their most popular model). Google has announced that Gemini is processing 746T per week; Claude is probably processing more. And the Chinese…
DeepSeek's official API, which has 10x cheaper cached input cost isn't even on OpenRouter as a provider, so just like Google, most volume is not going through OpenRouter. (Gemini's official hosted api is on OpenRouter BTW)
Also you're comparing an API with Google's internal corporate and consumer app use. Bytedance announced they were using 63T tokens/day (441T / week) at the end of 2025, so they are probably even higher than Google now. We don't know how much weekly tokens the DeepSeek chatapp uses, but it would also be a very high number much higher than OpenRouter tokens.
For the real reason of the recent price drops, go ask your AI about how much it would cost to run DeepSeek V4 or MiMo 2.5 after Ascend 950 PR have started to be mass delivered in 2026 Apr at $10k / card.
Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
#119Hot take: The reason this is happening is because the market for Chinese AI models hosted by Chinese companies is struggling. Even the market for Chinese AI models hosted by western companies is soft: During the week of May 18, OpenRouter processed 3.4T DeepSeek v3 Flash tokens (their most popular model). Google has announced that Gemini is processing 746T per week; Claude is probably processing more. And the Chinese…
I mean, I am going to use the best I can afford. And at work that's Opus, but while work is happy to let me spend $50+/day, that's just not viable for personal hobby use, I need to keep that in the realm of a WOW/mmo subscription.
Re: Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
#120Earlier quoted context omitted.
LLM providers can't "capture" anything. People loved Claude Code because it was cheap and good. Not cheap anymore? People switching to Codex, DS4 etc. Their only moat is maybe being SOTA but that only lasts so long before everyone else catches up.
I mean there is a minor moat. Most people don't enjoy switching providers or models. If you can get people to trust you'll stay near frontier, they'll stick around even when you aren't the best. Claude is a prime example of this
/model in OpenCode
There is no "moat" for me. Using the standard chat applications as a normal conversational/question has a little bit of moat as its able to cross reference existing conversations, but I disable that mostly anyways to prevent as much data retention as possible.