I used their Coding Plan for a few months. It is genuinely difficult to keep up with the models. The output is so fast. Qwen 3.8 27B is likely one of the strongest models they've hosted so far. Edit: it looks like this is only available on a API token pricing. Does anyone know if they have rolled out prompt caching yet? It used to get pretty expensive for agentic coding tasks with no prompt caching.
Qwen 3.8 27B available on Cerebras at 1500 tokens/s
11–20 of 237 posts
Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s
#12used their Code product with GLM4.7. its fun but if the model is bad it just doesn’t do much useful. Hope they add such models to Code too :)
Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s
#13Psychopaths: tok/SEC
Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s
#14For those who haven't noticed though, the context size they allow for Qwen is just 128k. Still interesting as a specialized sub-agent but not really well suited for long tasks.
Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s
#15Normal people: tok/s or t/s Psychopaths: tok/SEC
Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s
#16They do appear to host other models on OpenRouter so maybe Qwen3.8 will be there soon: https://openrouter.ai/provider/cerebras
Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s
#17Why do they only host small models rather than the 2.4T version? Is the I/O and interconnect between the wafers bad due to the limited beachfront relative to the massive size of the chip?
Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s
#18Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s
#19Normal people: tok/s or t/s Psychopaths: tok/SEC
Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s
#20Why do they only host small models rather than the 2.4T version? Is the I/O and interconnect between the wafers bad due to the limited beachfront relative to the massive size of the chip?
The wafer only has space for 44 gb of sram. If they offload ram they lose the speedup of having everything on 1 chip (the whole point of cerebras).
[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...