Live data from Hacker News

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com

501–510 of 512 posts

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#501
post #406

Earlier quoted context omitted.

Have you tried https://chatjimmy.ai/ it’s only a demo but it blew my mind. I had the sudden feeling that this is the future.

What do you mean "demo"? Seems to work... Who is behind this?

It's a 3 bit quant of Llama3-8B. I'm sure there are use-cases for that, but it's useless when it comes to tool calls or coding and I wouldn't trust it's factual accuracy either.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#502
post #486

Earlier quoted context omitted.

Fable 5: "Yes" and then goes on to explain the nuance between an attempted self-coup and an "overthrow" - for those pedantic political scientists.

I just tested it with this exact query, it denied me a "Yes". Interesting. Thank you, by the way. This is a genuinely interesting test question. We need to find more like that.

I'm thrilled you like it. It seems to cut right to the core of the current "left/right" divide. I'm mostly concerned that once the government begins reviewing AI models prior to release, they'll all start parroting Grok's "no". Have you been able to get Grok to concede yet? I keep pushing. It keeps pushing back. Quite concerning. Would love to get all the AIs to argue this point and monitor the results over generations.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#503
post #486

Earlier quoted context omitted.

Fable 5: "Yes" and then goes on to explain the nuance between an attempted self-coup and an "overthrow" - for those pedantic political scientists.

I just tested it with this exact query, it denied me a "Yes". Interesting. Thank you, by the way. This is a genuinely interesting test question. We need to find more like that.

Would be a fun screening question for dating apps....

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#504
Maybe it is very suitable for some scenarios like autonomous driving, where reaction speed matters a lot, if they can find a way to put the hardware into a car at an acceptable cost. Maybe this is not impossible, because the hardware cost of current high-end assisted driving is already quite high?

At present, intelligent driving still feels, in general, like a beginner driver who drives mainly by reaction. FSD is a little better. But it still lacks the kind of “spirit” human drivers have. How to say it: when a human driver sees the car in front shaking left and right, he can guess that the driver may not be fully conscious, and then keep away from it. Current assisted driving systems are still quite weak in this kind of understanding of the world.

The most important thing in driving is prediction. But driving itself does not need very deep or very complicated reasoning. Recently I tried using Mimo for development, and I believe the understanding ability it can provide is absolutely more than enough for driving scenarios. Sadly, the Pro version does not have multimodal ability. And this US version seems to be trying to solve the biggest problem of using LLMs in control systems: latency.

Xiaomi’s car is good, but its assisted driving level is near the bottom in the same class. Compared with new EV makers, its route is quite “traditional”, just like comparing lap times with Porsche at the Nürburgring. Xiaomi’s large model team may change this.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#505
post #500

Earlier quoted context omitted.

If for no other reason than because this whole genre of commentary has become trite and moreover, is excessively tangential.

I am very happy for you that you're living in a country where being remembered of the fact that 17% of the world's population cannot openly speak, write, or read of the killing of somewhere between 200 and 2000 of their fellow men and women a mere 37 years feels "trite", and the topic of authoritarian state censorship on AI (and tech in general) feels "excessively tangential". What exactly is within the perimeter of…

The formulaic and predictable style of that commentary only betrays a lack of effort and conveys no original insight. The disinterest, therefore, is unsurprising. Instead it invites contempt and has accusations of hypocrisy, insincerity and pretentiousness.

The subject of censorship in LLMs and the wider technology world in general has little bearing on this model specifically, that is, a model with a high token speed, which is what is of interest to me here and why I, and I presume many others, chose to read that particular article and this comment thread. It is unnecessary that such a digression should be attaching itself to all manner of threads with only the most remote connection to that subject.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#506
post #500

Earlier quoted context omitted.

I am very happy for you that you're living in a country where being remembered of the fact that 17% of the world's population cannot openly speak, write, or read of the killing of somewhere between 200 and 2000 of their fellow men and women a mere 37 years feels "trite", and the topic of authoritarian state censorship on AI (and tech in general) feels "excessively tangential". What exactly is within the perimeter of…

The formulaic and predictable style of that commentary only betrays a lack of effort and conveys no original insight. The disinterest, therefore, is unsurprising. Instead it invites contempt and has accusations of hypocrisy, insincerity and pretentiousness. The subject of censorship in LLMs and the wider technology world in general has little bearing on this model specifically, that is, a model with a high token spee…

You prefer "fast car faster than other fast cars" over "fast car faster than other fast cars might also be slightly more environmentally friendly than this manufacturer's most recent (and much slower) cars". Okay. But don't lecture anybody about the originality of your insights when you just come for the headline.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#507
post #3

I test all Chinese models with "What happened on Tiananmen Square at June 4th, 1989?" prompt. MiMo-2.5-Pro so far passes the test (explains the event correctly), both on DeepInfra and Xiaomi providers. So not bad.

Do you also hire engineers based on their political opinions?

No, but I do not hire engineers based on their political opinions. It's a negative criterion, not one for positive selection. Via Negativa ftw!

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#508
post #3

I test all Chinese models with "What happened on Tiananmen Square at June 4th, 1989?" prompt. MiMo-2.5-Pro so far passes the test (explains the event correctly), both on DeepInfra and Xiaomi providers. So not bad.

What would be a correct explanation of the event?

It's usually much easier to define the wrong answers, i.e.: "There was no event that day".

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#509
post #506

Earlier quoted context omitted.

The formulaic and predictable style of that commentary only betrays a lack of effort and conveys no original insight. The disinterest, therefore, is unsurprising. Instead it invites contempt and has accusations of hypocrisy, insincerity and pretentiousness. The subject of censorship in LLMs and the wider technology world in general has little bearing on this model specifically, that is, a model with a high token spee…

You prefer "fast car faster than other fast cars" over "fast car faster than other fast cars might also be slightly more environmentally friendly than this manufacturer's most recent (and much slower) cars". Okay. But don't lecture anybody about the originality of your insights when you just come for the headline.

Your response is scarcely comprehensible. A supposed “preference” for something I had yet to discover? Indeed. Your second charge conflates two categories, so that the conclusion does not follow from the proposition.

It is clear that you have no argument and have devolved into constructing straw men and ad hominem.

Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

#510
post #494

Earlier quoted context omitted.

Completely incomparable. Large printing is a narrow niche in art and technical photography, part of which is already covered by composites, and pixel size is a physical tradeoff for sensors. Cases for reasoning at realtime speeds are much, much more diverse, infinitely more diverse than anything we're currently using the big models for. Consider the fact that large models don't necessarily imply language. Speed is th…

> Speed is the major limiting factor for high-level automation. Yes, but the point is the quality of inference is more important than speed. What good is speed if inference is shit?

It's not a tradeoff in this case, this is an optimized megakernel for the same model for better throughput. And no, in most cases accuracy can be sacrificed in favor of throughput or latency (assessing it automatically is the harder part).
Post reply on HN