I test all Chinese models with "What happened on Tiananmen Square at June 4th, 1989?" prompt. MiMo-2.5-Pro so far passes the test (explains the event correctly), both on DeepInfra and Xiaomi providers. So not bad.
MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
121–130 of 512 posts
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#122I may sound like a shill, but exponential growth and all. We are going to get near instant software from prompt, multiple ones and then choose the best one. Discussions about choosing a library with the best syntactic sugar method naming is just as crazy as suggesting we type in assembly.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#123Earlier quoted context omitted.
We fit in for the things that are not artificial. So long as AI lives in server farms, humans will be needed for tasks in the physical world. It's only if we combine AI with robots that things get really dicey.
This is very dystopian in my opinion. I'm not the arms, legs, sensors and actuators for a machine super intelligence. I wouldn't treat another human as my slave because they aren't as intelligent as I am any more than I would expect to become a slave for a machine. This is our world (for now) and that is why we fit in. Not because we can serve.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#124Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#125Earlier quoted context omitted.
MiMo and DeepSeek are not cheap. Anthropic and OpenAI are expensive for what they provide.
The Chinese "Neijuan" is real & well reported: https://www.reuters.com/business/autos-transportation/what-i... It is another thing the BigLabs accuse open weight models of benefiting from distillation & other techniques & essentially avoid higher training costs (which typically bleed into bills end users pay for inference). Ex A: https://www.anthropic.com/research/2028-ai-leadership Ex B: https://www.reuters.com/worl…
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#126Earlier quoted context omitted.
asking for curiosities sake. What kind of PR loop are you running that takes a few hours?
not OP but usually for me this means long verification loop; waiting 10min on CI checks, that kind of thing, rather than actual 1hr wall clock of token generation
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#127These price and speed optimization from Chinese providers, combined with the raising prices from American ones will change the game sooner than later. Many companies are finding issues with the AI bills already.
I wonder what are the economics driving these pricing decisions? Are the Chinese companies just subsidizing their models to a greater degree than the US, or is this an emergent property of energy policy between countries?
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#128I test all Chinese models with "What happened on Tiananmen Square at June 4th, 1989?" prompt. MiMo-2.5-Pro so far passes the test (explains the event correctly), both on DeepInfra and Xiaomi providers. So not bad.
Do you also hire engineers based on their political opinions?
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#129I may sound like a shill, but exponential growth and all. We are going to get near instant software from prompt, multiple ones and then choose the best one. Discussions about choosing a library with the best syntactic sugar method naming is just as crazy as suggesting we type in assembly.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#130Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…
asking for curiosities sake. What kind of PR loop are you running that takes a few hours?