So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…
MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
251–260 of 512 posts
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#252Earlier quoted context omitted.
It's going to be an issue when China ends up scaling faster as well. Faster tokens, faster clusters, qat models, fp4, it's getting scary.
Issue for who?
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#253It will be cool to measure models based on their RAW performance and measure them in terms of ROI - not some benchmark but something meaningful like we used this model to solve X.
That will be a massive mind shift and might justify the token expenditure.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#254Earlier quoted context omitted.
I'm using Deepseek-v4-pro as my main model and this is sometimes pretty annoying, I have to do some easy boring task, think "I'll just leave the agent to do it and go take a nap", but it's already done writing the code before I even walk away from the computer
Do you mean Flash and not Pro? I haven't tried it personally, but according to OpenRouter, the fastest DeekSeep V4 Pro providers are only ~50tps. That's slower than Claude Opus. https://openrouter.ai/deepseek/deepseek-v4-pro?sort=throughp...
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#255How? edit: now I read the article fully, seems like they utilize some very effective MTP algorithm. and somehow the quality is still decent enough. though, I doubt that the quality really only drip a bit like they claimed. maybe for the benchmarks, but for general uses the heavily quantized models very often so worse result.
They say they are using https://github.com/tile-ai/TileRT - persistent CUDA kernel - tiled processing with overlapping read/writes - model designed with specific constraints in mind
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#256Earlier quoted context omitted.
MiMo and DeepSeek are not cheap. Anthropic and OpenAI are expensive for what they provide.
You don't consider Input $0.435 Output $0.87 cache read $0.003625 per million tokens for near frontier intelligence cheap?
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#257Earlier quoted context omitted.
>It's such a weird "Gotcha" that seems to only assume that Chinese LLMs might censor something. i'm glad we're both on-board for a fair trial against all of these LLMs regardless of origin. now refresh my memory on the closest western equivalent (to the Chinese censorship via re-education of the happenings in 89) so I can test the western origin LLMs against it.
the civil war was only ever and exclusively about states rights
> The U.S. Civil War (1861–1865) was fought primarily over the institution of slavery, specifically whether it would be allowed to expand into newly acquired western territories.
> While you might hear people point to "states' rights" or economic differences as the causes, these issues were inextricably linked to slavery. The southern states wanted the "right" to maintain and expand slavery, while the northern states increasingly opposed its expansion.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#258Fast AI seems genuinely exciting and somewhat unsettling to me. Right now Claude is faster than me on some tasks but we’re at least close. I have a prompt to clean up a PR that’s been running for 1h now and I expect it to take another few. It’s hard to imagine how the workflow would look like if it was near-instant. On the one hand, it might be easier to focus. Some prompts take so long that I start to multitask and…
But truly, using Cerebras at ~2k tokens/s, with very low latency is like a vision into the future. You start to rework your workflow around things that can happen without onerous manual review - stating the conditions for success, etc. It's rare that I have a problem that maps well to that, but I expect this is where things are headed.
Of course the fast models tend to not be the SOTA ones, but if that was the case - high quality and near-instant thinking, that's a game changer that I don't think we're really prepared for. The things that get unlocked with higher-than-reasonable speed become very interesting.
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#259Speed is indeed a next big thing what should happen with LLM frontier models. The possibilities with current models but 1000 times faster would be super useful. Earlier this week it took Claude at least full time a week with two max subscriptions to solve a complex issue where we wanted to mimic a occlusion mapping variant used in the game Crimson Desert. Pretty complex mathematical challenge. With a ultra fast LLM a…
Re: MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
#260So, regarding the productivity argument: I don't get it. It doesn't really matter (for regular employees) that you can do now in 2h what before it took 2 days. Why? Because it's not that you have the rest of the day for yourself. You still have to work 8h/day as usual. But now the pattern is different: instead of enjoying the craft digging deeper into problems in the span of 2 days, now you are rushing into some slot…
Sure but if you're really unhappy with your employer employeeing you for 8 hours a day you can also harness this power on your own personal projects to help break free from the 9-5 grind if you so desire.