Live data from Hacker News

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

inference-docs.cerebras.ai

101–110 of 239 posts

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#101

150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more flexible rate pool. Even trying it out, it seems like our account has gotten moved to some limbo where we can no longer add billing information. ``` Billing access restricted Self-serve billi…

What kind of coding tasks would you expect to hit that limit? In my setup, on a very large codebase, it takes each agent 3-4 minutes at minimum to go past 100k tokens.

(note it's 150k uncached tokens, the total limit is 450k/min)

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#102
post #65

Earlier quoted context omitted.

The target audience is who needs raw speed. Having the choice is good as you can make a trade-off between speed, perf, and quality. Until last year, people had a single AI god they believed in (mostly Anthropic stuff). Now we have power to make choices (open-weights, SOTA, speed-optimized, etc) the same way you do for system designs.

It's not a criticism, I was really looking forward to trying out such a powerful model at this speed. But I burn my 5$ allowance in 10 minutes ... and only because I was hitting rate limits, without it would probably be less than a minute.

I hear ya... the best option is to use company budget as normies will rack up ridciulous amount soon with that raw speed.

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#103

150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more flexible rate pool. Even trying it out, it seems like our account has gotten moved to some limbo where we can no longer add billing information. ``` Billing access restricted Self-serve billi…

What kind of coding tasks would you expect to hit that limit? In my setup, on a very large codebase, it takes each agent 3-4 minutes at minimum to go past 100k tokens. (note it's 150k uncached tokens , the total limit is 450k/min)

in my last tests with cerebras for coding tasks, most large tasks or anything greenfield would hit token limits. note that smaller models and the gpt-oss-120b style models they used to run are very prone to overthinking, so individual turns may be 3-10k tokens of just thinking + input + output.

i don't think it's quite apples-to-apples to compare to a frontier model or even a k3. the odds of success (file compiles? read the right context?) are lower and thinking is longer.

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#105
The question is whether Cerebras is available... I've been trying to get https://www.cerebras.ai/code for at least 1 year now. It's all sold out. Always. I once joined their Discord, waited for the drop, and it all sold out in seconds. I haven't had enough time to put my card details. Somebody recommended that I should put my card details in advance, lol.

The next time I hear about them I am laughing, because when I could enjoy these powers? How many years I should be sitting in a waitlist...

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#106
Do I understand their pricing correctly? This is $10 per month for a developer account PLUS you pay $1.49/M for output tokens and $0.99/M for input tokens on Qwen 3.8 27b with a 128k context?

EDIT: Or, maybe it's just token pricing, but $10 is the minimum? Maybe it's that.

https://www.cerebras.ai/pricing

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#107

Do I understand their pricing correctly? This is $10 per month for a developer account PLUS you pay $1.49/M for output tokens and $0.99/M for input tokens on Qwen 3.8 27b with a 128k context? EDIT: Or, maybe it's just token pricing, but $10 is the minimum? Maybe it's that. https://www.cerebras.ai/pricing

No. You buy a minimum of $10 worth of credit, then use it at $1.49/M rate. There is no recurring charge.

There is a separate subscription based plan, which is sold out now.

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#108
post #95
post #92

Earlier quoted context omitted.

This was my experience a year ago on some other model they could run super fast. Routine coding tasks would hit the per-minute token limits. Just the math there... 150k TPM... and 15k TPS means... you can run for 10 seconds every minute? The basic math boggles the mind.

Not sure how the rate limiting works, but it's 1.5k TPS, not 15k, so you could run it for 100s/min, which seems good enough to me

It seems you forgot to account for the fact that cerebras uses a baker's minute which is 144 seconds instead of 60. (Seriously though what's the supposed issue here?)

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#109
post #95

Earlier quoted context omitted.

Not sure how the rate limiting works, but it's 1.5k TPS, not 15k, so you could run it for 100s/min, which seems good enough to me

iirc input (uncached) goes towards the limit as well

What's the tok/s when they process input?

Re: Qwen 3.8 27B available on Cerebras at 1500 tokens/s

#110

Qwen 3.8 27B is an exceptional model for coding and ranks as one of the best local models for coding....BUT in my head I am confused why a company that's IPO'd doesn't invest in RL'd super specialized, super-damn-fast models for very specific tasks - instead of giving us the OSS GPT model from what feels like 200 years ago

almost of their business is hosting Sol ultra fast or whatever for OpenAI to use internally
Post reply on HN