Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

301–310 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#301

Earlier quoted context omitted.

In which world do you live where employees work 8 hours per day ? They clock 8 hours per day maybe, but they don't work that time

Some companies force you to actually work 8 hours a day. It’s hell.

Which country and which companies ?

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#302
post #285

Earlier quoted context omitted.

In which world do you live where employees work 8 hours per day ? They clock 8 hours per day maybe, but they don't work that time

In theory, ofc. But that doesn't matter. If you were doing something that took 2 days in average, but you were doing it in half the time, then that was fine pre LLMs. Nowadays your manager knows that with LLMs you need to deliver faster no matter what, and then it's more difficult to "hide" and to slack.

Yeah. So, good things. We ack know that people are mostly slacking at work

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#303

Earlier quoted context omitted.

Agent mania setting in It's also pretty funny sometimes how it gives weird future roadmap estimates ("part 2 - 3 weeks, part 3 - 2 months", etc.) and when you tell it to actually do those changes it's pretty much done in half an hour

I've long believed those numbers were faked by Anthropic/OpenAI to serve as a form of advertisement. The estimates are impossible to verify and their ability to do "2 days of work" in 10 minutes will presumably make the user go "Wow, I just saved SO much time!" Plus, the unnecessary text eats up the users' tokens so it helps the companies on the backend, as well.

Nah it’s all from the pretraining data

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#304

Earlier quoted context omitted.

You don't consider Input $0.435 Output $0.87 cache read $0.003625 per million tokens for near frontier intelligence cheap?

No. They still have enormous profit margins on inference with these prices.

Their margins doesn't impact my own assessment of end user pricing as cheap.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#305

Below is the part I found most interesting > "However, naively applying FP4 across the entire model causes degradation in complex reasoning, logic, and code generation. Given the MoE (Mixture of Experts) architecture of Xiaomi MiMo-V2.5-Pro — where Experts constitute the vast majority of parameters and exhibit the highest tolerance to quantization — we selectively quantize only the MoE Experts to FP4 while preserving…

The 120B and 20B GPT-OSS models by OpenAI did this last year for what it’s worth; the MoEs where MXFP4

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#307

Earlier quoted context omitted.

Odd, I'm having the opposite experience. The thing I really love about working with computers is when I achieve something. That's the thing that makes me figuratively, and sometimes literally, throw my fists into the air and go "Yeaaah!" With the AI tooling, I'm getting those more like a couple times a week. Plus, I'm using AI to attack the things in my day that are "a drag", and getting them done too. The highs are…

I did a deep binge on two or three projects I would never do, and like five small ones that would have consumed months. It felt like that, kinda, for a bit. Now whenever it does something for me I get nothing. I didn’t do it… the chatbot did. What’s for me to celebrate? How can there be any real pride or satisfaction for a thing that was just handed to me because I asked for it? If anything it diminishes my satisfact…

I hear what you and the other sibling comment are saying. I, thankfully, somehow, am able to focus more on the results than the process. Having fun playing a game (that AFAIK no longer exists) with my family is still having fun. Having people using a new apt cacher that fixes problems with existing ones, and also can survive the recent DDoS, is still a really great thing.

But, I'm not going to yuck your yum. I appreciate the people who do jointery using hand tools, even if I'm out here with a track saw and a router.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#308
post #120

Earlier quoted context omitted.

Energy is likely more abundant in China. I am not sure about compute, but that must be part of reason for such drastic price differences.

They're leaving us in the dust on solar, while our current administration is still trying to put people in the ground to dig up more coal and die of black lung. https://en.wikipedia.org/wiki/Solar_power_in_China

They're building more coal than anyone.

Also more nuclear than anyone, which one must assume you hate, because preferring solar requires you don't actually understand thing

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#309
post #76
post #17

These price and speed optimization from Chinese providers, combined with the raising prices from American ones will change the game sooner than later. Many companies are finding issues with the AI bills already.

Another problem is that US models are all closed source, and if you're a large corporate you may not want your org to be held hostage by OpenAI / Anthropic. I genuinely don't understand what moat these US model labs have. If they're saying recursive self improvement is just around the corner and Chinese labs are only slightly behind the leading US models, what moat does the US labs have? Are the US models going to re…

maybe the moat is that we slowly start to forget how to code by hand and then you -need- the AI tool.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#310

Earlier quoted context omitted.

Agent mania setting in It's also pretty funny sometimes how it gives weird future roadmap estimates ("part 2 - 3 weeks, part 3 - 2 months", etc.) and when you tell it to actually do those changes it's pretty much done in half an hour

I've long believed those numbers were faked by Anthropic/OpenAI to serve as a form of advertisement. The estimates are impossible to verify and their ability to do "2 days of work" in 10 minutes will presumably make the user go "Wow, I just saved SO much time!" Plus, the unnecessary text eats up the users' tokens so it helps the companies on the backend, as well.

I tend to be cynical about AI companies, but I'm guessing the bad estimates more just come from a complete lack of actual data it could use for that so it's more or less a hallucination.
Post reply on HN