Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

321–330 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#321

Earlier quoted context omitted.

I’ve been playing around with groq and GPT OSS which they run at 1000 TPS (20B) or 800 TPS (120B) and the speed feels quite magical. I haven’t tried cerebras’ 3000 TPS yet but I did try the demo of that 15,000 TPS model whose name escapes me right now. I’m not sure if it makes a meaningful difference for my actual work, but it sure is amazing to watch it generate a screen full of text in the blink of an eye. I do thi…

https://chatjimmy.ai/ ?

AFAIK Taalas, the company behind this demo, still only have their initially "hardwarized" model available to test in ChatJimmy, which IIRC is a rather stupid Llama 3ish 8b.

Don't get me wrong though, that demo is still incredibly impressive & makes me very much excited for the hardware-based model era (potentially) ahead.

Once you've experienced those speeds, you really start to think about the whole class of things that becomes possible; massively parallel decode paths, extensive reasoning loops, etc…

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#322
post #19

Cerebras is trialing Kimi K2.6 at 3000t/s (invite only). I'm excited for when the fast hardware gets more mainstream for frontier models. Models designed for speed on Nvidia are nice addition that could bridge the gap.

Cerebras got lucky that they IPOed last month instead of now.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#323
post #58

Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…

[flagged]

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#324
post #136

So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…

In my case, I think slower model makes it hard to manage context and tasks in parallel. I would much prefer to work in one task only, and finish it, take a break, and work on another task. Currently I have three tabs for three tasks in parallel, it is much worse than because constantly context switching is painful. I think a faster model would mean that you don't have to start a new task while waiting.

Agents completing work faster would certainly help me as well since I also find context switching exhausting above some threshold.

Build and test would move back into the critical path, though, and for some projects that will take effort to bring down.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#325
post #321

Earlier quoted context omitted.

https://chatjimmy.ai/ ?

AFAIK Taalas, the company behind this demo, still only have their initially "hardwarized" model available to test in ChatJimmy, which IIRC is a rather stupid Llama 3ish 8b. Don't get me wrong though, that demo is still incredibly impressive & makes me very much excited for the hardware-based model era (potentially) ahead. Once you've experienced those speeds, you really start to think about the whole class of things…

For scale though if three or four chips that size can replicate a Qwen 27B experience that'll be quite useful.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#326

Earlier quoted context omitted.

I think people are continuing to view these systems as pure LLMs - when that ship sailed 6+ months ago. Between being able to review memory, using agent harnesses and sub agents and skills to go out and discover information - modern systems (Codex, Claude Code, Cursor) - use LLMs - but the LLM is only a small component of it. Compare what you get from sending a request to a chatbot like ChatGPT - to what you can from…

You think someone is, or even should, special case things like estimates? What else deserves that level of intervention so they look less dumb? Logistics for getting to the car wash next door? In the mean time, alas, no, we can see from actual prompts sent directly or through sub-agents, and actual replies, estimates remain LLM generated. Though, this discussion here could change that, because indeed there is a lot o…

rather than special casing, make real data based on chat logs for how long things took both in calendar and chat time

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#327
post #136

So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…

You can dig deeper into problems with AI. For me, it supplements my knowledge in domains I don’t fully understand. It also helps me learn. So I can tackle problems I wouldn’t otherwise. I’m excited for ultrafast AI. It likely means less temptation to multi-thread and deeper flow in single sessions.

how do you know that it is actually suggesting the right thing?

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#328
post #136

So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…

In which world do you live where employees work 8 hours per day ? They clock 8 hours per day maybe, but they don't work that time

Here’s my hot take as an elder millennial. Boomers are the absolute worst at being unable to make the distinction between time at work and time doing work. They may show up an hour before everyone else but spend the first two or three hours a day, reading the news and getting coffee and making small talk and accomplishing literally nothing. Then crow about their work ethic.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#329
post #136

So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…

Like with any tech there are dumb ways of using it and there are smart ways. Treating it as a "slot machine giving you the right answer" is a dumb way - it may work for a bit, but it won't carry you very far because everyone else can also do this. No one is stopping anybody from digging deeper into problems than ever before using this technology - that's the smart way.

I'm amazed at how steep the AI learning curve continues to be and how people are spread so far apart on it. I think supercharged learning with AI and agents is undervalued at this point but that more people will realize its utility over time, especially as a complement to delegating work.

It also makes me think about the temptation to stop thinking with these tools, i.e. "cognitive surrender". Addy Osmani wrote a nice blog post about this: https://addyosmani.com/blog/cognitive-surrender

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#330
post #58

Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…

I’ve been playing around with groq and GPT OSS which they run at 1000 TPS (20B) or 800 TPS (120B) and the speed feels quite magical. I haven’t tried cerebras’ 3000 TPS yet but I did try the demo of that 15,000 TPS model whose name escapes me right now. I’m not sure if it makes a meaningful difference for my actual work, but it sure is amazing to watch it generate a screen full of text in the blink of an eye. I do thi…

> I haven’t tried cerebras’ 3000 TPS yet but I did try the demo of that 15,000 TPS model whose name escapes me right now.

You were likely thinking of AI accelerator startup Taalas.

Previous HN discussion: https://news.ycombinator.com/item?id=47086181

Post reply on HN