Live data from Hacker News

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

inference-docs.cerebras.ai

31–40 of 226 posts

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#31

It would be great if they made their inference capacity for this model available via OpenRouter; the fastest provider on OpenRouter right now is at ~80tps https://openrouter.ai/qwen/qwen3.8-27b#providers They do appear to host other models on OpenRouter so maybe Qwen3.8 will be there soon: https://openrouter.ai/provider/cerebras

We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).

https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#33
I have used their gemma 4 31b model through kagi and getting real instantaneous answers is absolutely crazy. A very different feeling and UX. Even if the model is smaller, there is definitely a use case for these. I was wondering if they would put the qwen 27b model, it sounds very interesting to try.

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#34
post #17

Earlier quoted context omitted.

The wafer only has space for 44 gb of sram. If they offload ram they lose the speedup of having everything on 1 chip (the whole point of cerebras).

They can host larger models by pipelining it on multiple wafers. Each wafer stores one layer and N layers can serve an N * 44 gb model with N concurrency. The limitation would of course be inter-wafer I/O, which my comment was getting at. That's probably how they can serve bigger models like GPT 5.6 Sol [1]. [1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...

I never said offloading was impossible. It will result in a large slowdown.

It would look bad for cerebras if other people are hosting the 27b version and show a higher TPS than cerebras.

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#35
post #2

I used their Coding Plan for a few months. It is genuinely difficult to keep up with the models. The output is so fast. Qwen 3.8 27B is likely one of the strongest models they've hosted so far. Edit: it looks like this is only available on a API token pricing. Does anyone know if they have rolled out prompt caching yet? It used to get pretty expensive for agentic coding tasks with no prompt caching.

It appears that they do support Prompt Caching: https://inference-docs.cerebras.ai/capabilities/prompt-cachi...

> How are cached tokens priced?

> There is no additional fee for using prompt caching. Input tokens, whether served from the cache or processed fresh, are billed at the standard input token rate for the respective model.

Well, talk about flipping the narrative.

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#36
Funnily enough the pricing isn't that much worse than on openrouter, where the best price at the moment is $0.24 in / $2.55 out, vs $1 / $1.5 on Cerebras.

Sure, 4x input , but cheaper output. Though Cerebras doesn't have prompt caching, so not great for agentic workloads. (they do, but it doesn't affect the price.

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#38
post #2

I used their Coding Plan for a few months. It is genuinely difficult to keep up with the models. The output is so fast. Qwen 3.8 27B is likely one of the strongest models they've hosted so far. Edit: it looks like this is only available on a API token pricing. Does anyone know if they have rolled out prompt caching yet? It used to get pretty expensive for agentic coding tasks with no prompt caching.

It appears that they do support Prompt Caching: https://inference-docs.cerebras.ai/capabilities/prompt-cachi...

It doesn't reduce the price though.

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#39

I really wish they had their customer support somewhere else than Discord, which seems to think I'm a bot and doesen't accept my email or phone numbe

discord support can fix such issues

If you need customer support to access customer support, something is wrong; no?
Post reply on HN