Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

401–410 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#401

Earlier quoted context omitted.

I see bigger problem with model inconsistency. You never know whether Anthropic will route your request to a cheaper model for the price of Opus. So you can never estimate how much a task will cost, because you might have to restart several times and pay for each attempt. Then you have to prompt models to gauge whether they are real or impostors which also adds to token usage.

> You never know whether Anthropic will route your request to a cheaper model for the price of Opus For non subsidized plans? Pretty sure they'd need to put this in ToS, or law suites would have followed by now.

1. How would you know?

2. They are doing lots of shady stuff that would have gotten someone else banned from visa/mastercard. Your paid off plan literally changes after billing...

I think people are letting them fly for now, because if it turns out true that they'll have AGI they want to be on their good side? We might see the knifes getting pulled otherwise.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#404
post #139
post #68

Earlier quoted context omitted.

We are living in a ZIRP-like era where builders at the fastest pace layer have misattributed their velocity to exponential gains in model capability. In fact, they are surfing on decades of careful effort to build a robust foundation of highly reusable software libraries. This strategy will seem to work really well until the economy that enabled that foundation to form is hollowed out. Then, there will be a reckoning…

This is a great point. LLMs can't speed up human decision processes and alignment.

Not entirely sure about that.

Its already speeding up human decision processes, and while ethics / alignment may seem unique to humans we also see normative expressions in monkeys or apes (like the experiment where one is given a grapes, the other cucumber).

A lot of ethics is based on symmetry: symmetric relations, equal rights, equal voting power, ... symmetries sound rather mathematical if you ask me, and decision structures have historically been pressed towards democracy (or at least depiction of it). One could say that modeling humanity as an empire with a king, ignores the will of sometimes hungry farmers with pitchforks. To prevent the occasional "implicit democracy" (royaltycide), it turned out in the interest of the king to recognize the powers of those farmers, and to formalize it in the decision making process. Or at least pretend to.

I believe machines will be able predict the preference sentient creatures would prefer in terms of decision structures, but I don't believe it will be able to predict (without human exposition) those novel preferences that stem not from sentience but from being specifically human properties (i.e. irritants which are quasi universal for humans, etc.), some of them humans know how to make predictions for (we can run expensive simulations modeling what happens when protein X is exposed to substance Y, and then make heuristic predictions of the effect on a full human in a realistic environment). So at a fundamental level I agree: machine learning models are not guaranteed to help much in predictions concerning entirely unexplored territory, neither by humans nor by natural selection. But it will definitely be capable of replacing the average human job, which doesn't involve consensual exploration outside of the homeostasis required in the implicit job description, that seems entirely automatable, regardless if its physics, mathematics, (harder than computer science), let alone programming.

It won't be able to magically systematically correctly predict out of distribution datapoints, it could only explore it like humans could by trial and error.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#405

Earlier quoted context omitted.

The sanctions only “prevent” them from directly buying NVidia’s latest and greatest in the sense that NVidia can’t sell directly to them. Essentially, there are companies now who are in a country without the sanctions, they buy from NVidia (or a partner), and then ship them off to China. For the orgs in China doing this, there’s zero legal risk besides having foreign customs service intercept the shipment and losing…

That wouldn't explain why Deepseek is fast relative to other Chinese providers, especially considering that they're reportedly ahead of the curve among Chinese companies in moving off Nvidia. I think their quant fund background has more to do with it. Their models are clearly designed with performant inference clearly in mind.

Yes, it's performant, and esp performant at non-trivial context depths. DeepSeek-V4 DS4 (and Flash - DS4F) drop tok/s speed much less than the rest. On my M2 Max it took context depths of 768K to drop tok/s to ~10 tok/s.

https://x.com/ljupc0/status/2062457314414587996

Other local models I've checked drop to unusable speeds way sooner. Only other model with similarity favourable curve I've tried is nemotron-cascade-2-30b-a3b. But it's a small model, way dumber than DS4F.

Coding agents use cases have large context depths. The rate of decline is as important as the headline number.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#406
post #58

Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…

Now the next bottleneck is the compiler - which we can model in an LLM! It's only wrong 15% of the time :) But truly, using Cerebras at ~2k tokens/s, with very low latency is like a vision into the future. You start to rework your workflow around things that can happen without onerous manual review - stating the conditions for success, etc. It's rare that I have a problem that maps well to that, but I expect this is…

Have you tried https://chatjimmy.ai/ it’s only a demo but it blew my mind. I had the sudden feeling that this is the future.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#407
post #16

I may sound like a shill, but exponential growth and all. We are going to get near instant software from prompt, multiple ones and then choose the best one. Discussions about choosing a library with the best syntactic sugar method naming is just as crazy as suggesting we type in assembly.

How do you get all the build system scripts/tests.... to run instantly?

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#408
post #334

Earlier quoted context omitted.

Actually in this case they possibly are estimates. It's been known for some years[1] that LLMs do regression in-context. Frontier models have been trained against many, many issue text that include task break downs and estimates. [1] https://arxiv.org/html/2409.04318v1

Interesting. So it may have learned how to estimate as a human but doesn’t understand that it doesn’t operate at that speed :D I wonder if there’s a reasonable way to give an llm parameters that give it a concept of its own execution speed. Seems that could be useful for multiple purposes

Yes, it's entirely possible to do that via RL. It'd be a fun little project you could do for less than $100 on a small LLM actually.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#409

Earlier quoted context omitted.

I did a deep binge on two or three projects I would never do, and like five small ones that would have consumed months. It felt like that, kinda, for a bit. Now whenever it does something for me I get nothing. I didn’t do it… the chatbot did. What’s for me to celebrate? How can there be any real pride or satisfaction for a thing that was just handed to me because I asked for it? If anything it diminishes my satisfact…

> The things I had to learn and the informed decisions I had to make? All pointless trivia, now. A child could do it. Probably this is a hyperbole. Did you do the experiment? I expect that the child won't be able to do it. Ask an adult. Same thing. Ask an expert of the domain. Maybe but not as fast or as good as you.

Yes that’s more “how it feels” than something I’ve had kids actually try.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#410
The interesting bits on how they achieved it:

> On the model side, we applied FP4 quantization

> introduced DFlash, an efficient speculative decoding method based on block-level masked parallel prediction

> On the system side, TileRT perfectly adapts to the dynamic characteristics of these algorithms

> 1000+ tokens/s output [...] using just a single standard 8-GPU commodity node

Post reply on HN