Wow. So kimi is suddenly half the price for all users until they hit 256k of context? Thats massive.
Kimi K3-256k
121–130 of 172 posts
Re: Kimi K3-256k
#122Re: Kimi K3-256k
#123i've never had any issues with 256k context. see no reason to bump up to 1m if it comes at a premium.
Re: Kimi K3-256k
#124Re: Kimi K3-256k
#125I was so excited that it is open source until i realised the model required 1.5To of VRAM. Unsloth has compressed it in 1bit at about 570gb VRAM with 75% accuracy, that's almost mac studio territory...
Re: Kimi K3-256k
#126Earlier quoted context omitted.
Not surprising it's a hard cutoff: they almost certainly have two infrastructure configurations for the two max sequence lengths Fewer nodes dedicated to prefill per instance, and fewer nodes in total since you don't need to support a higher KV cache. Disaggregated inference also means they can tune the balance of compute dedicated to prefill seperately from decode
AI infra buildup is so massive that the frontier labs should be able to offer more than one level of context length to incentivize token thriftiness. One would think compute-constrained actors like Anthropic would have done so, unless prefill isn’t really a bottleneck compared to decode?
Re: Kimi K3-256k
#127Re: Kimi K3-256k
#128I was so excited that it is open source until i realised the model required 1.5To of VRAM. Unsloth has compressed it in 1bit at about 570gb VRAM with 75% accuracy, that's almost mac studio territory...
Re: Kimi K3-256k
#129[flagged]
Re: Kimi K3-256k
#130i've never had any issues with 256k context. see no reason to bump up to 1m if it comes at a premium.
It's pretty handy for have very long contexts for long running agents, or else when doing literary analysis to simply be able to load the entire book in.