It's interesting how even 5 tok/s is still much faster than you'd typically type, but feels glacially slow for an agent. On the other hand, I've been using Mimo and Minimax a lot recently. They routinely reach 100-150 tokens per second and that feels too fast , to the point where it's hard to keep up with what it's actually doing. Great for subagents though.
> It's interesting how even 5 tok/s is still much faster than you'd typically type, but feels glacially slow for an agent. Calling the token rate the rate at which they "type" is a bit misleading. They also do virtually all of their more complex reasoning in tokens, so 5 tokens per second is also their thinking speed. And thinking at 5 tokens per second is glacially slow. This is why faster versions of strong models…
How fast is N tokens per second really?
91–100 of 105 posts
Re: How fast is N tokens per second really?
#92Earlier quoted context omitted.
> It's interesting how even 5 tok/s is still much faster than you'd typically type, but feels glacially slow for an agent. Calling the token rate the rate at which they "type" is a bit misleading. They also do virtually all of their more complex reasoning in tokens, so 5 tokens per second is also their thinking speed. And thinking at 5 tokens per second is glacially slow. This is why faster versions of strong models…
How many tok/s does an average human think?
Re: How fast is N tokens per second really?
#93I wonder when we reach speed of 1000 tps with high quality models. 5 years? 10 years?
Maybe when intelligence plateaus it could become a main differentiating factor, like smartphones and battery life.
Re: How fast is N tokens per second really?
#94I wonder when we reach speed of 1000 tps with high quality models. 5 years? 10 years?
Since the whole goal of software architecture schemes it to allow the rest of us non-geniuses to still understand it and modify it, perhaps the same could be true of llms.
Perhaps a million-per-second hypothetical (small) model can be more useful than a state of the art big one.
Re: How fast is N tokens per second really?
#95Something like https://chatjimmy.ai/ will do 14,000 tokens per second, and it's a completely different experience to what we see now. It's more like a page load than a conversation.
Re: How fast is N tokens per second really?
#96I always wonder what the world is going to be like when these are 2 orders of magnitude faster. Something like https://chatjimmy.ai/ will do 14,000 tokens per second, and it's a completely different experience to what we see now. It's more like a page load than a conversation.
I'd much rather trade speed for accuracy.
Re: How fast is N tokens per second really?
#97I always wonder what the world is going to be like when these are 2 orders of magnitude faster. Something like https://chatjimmy.ai/ will do 14,000 tokens per second, and it's a completely different experience to what we see now. It's more like a page load than a conversation.
That's really impressive from a speed point of view but it hallucinates like it's on drugs. I'd much rather trade speed for accuracy.
Re: How fast is N tokens per second really?
#98Re: How fast is N tokens per second really?
#99Earlier quoted context omitted.
> It's interesting how even 5 tok/s is still much faster than you'd typically type, but feels glacially slow for an agent. Calling the token rate the rate at which they "type" is a bit misleading. They also do virtually all of their more complex reasoning in tokens, so 5 tokens per second is also their thinking speed. And thinking at 5 tokens per second is glacially slow. This is why faster versions of strong models…
How many tok/s does an average human think?
Re: How fast is N tokens per second really?
#100One can say the most profound thing in 3 words, slowly or fast it does not even matter, and can also spout absolute senseless garbage in billions of words at absolutely ridiculous speed.