Live data from Hacker News

Cerebras launches Qwen3-235B, achieving 1.5k tokens per second

cerebras.ai

91–100 of 160 posts

Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second

#91
post #18

Cerebras is truly one of the maddest technical accomplishments that Silicon Valley has produced in the last decade or so. I met Andy seven or eight years ago and I thought they must have been smoking something - a dinner plate sized chip with six tons of clamping force? They made it real, and in retrospect what they did was incredibly prescient

The concept is super cool but does anyone actually use them instead of just buying Nvidia?

Mistral uses them for Le Chat, it is really fast.

https://chat.mistral.ai/chat

https://www.cerebras.ai/blog/mistral-le-chat

Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second

#92

If this is the full fp16 quant, you'd need 2TB of memory to use with the full 131k context. With 44GB of SRAM per Cerebras chip, you'd need 45 chips chained together. $3m per chip. $135m total to run this. For comparison, you can buy a DGX B200 with 8x B200 Blackwell chips and 1.4TB of memory for around $500k. Two systems would give you 2.8TB memory which is enough for this. So $1m vs $135m to run this model. It's no…

>Maybe hedge funds or some sort of financial markets?

I'd think that HFT is already mature and doesn't really benefit from this type of model.

Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second

#94
post #92

If this is the full fp16 quant, you'd need 2TB of memory to use with the full 131k context. With 44GB of SRAM per Cerebras chip, you'd need 45 chips chained together. $3m per chip. $135m total to run this. For comparison, you can buy a DGX B200 with 8x B200 Blackwell chips and 1.4TB of memory for around $500k. Two systems would give you 2.8TB memory which is enough for this. So $1m vs $135m to run this model. It's no…

>Maybe hedge funds or some sort of financial markets? I'd think that HFT is already mature and doesn't really benefit from this type of model.

True, but if the hardware could be “misused” for HFT, it’d be awesome.

Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second

#96
post #13
post #11

With this kind of speed you could build a large thinking stage into every response. What kind of improvement could you expect in benchmarks from having say 1000 tokens of thinking for every response?

Thinking can also make the responses worse; AIs don't "overthink", instead they start throwing away constraints and convincing themselves of things that are tangential or opposite to the task. I've often observed thinking/reasoning to cause models to completely disregard important constraints, because they essentially can act as conversational turns.

> start throwing away constraints and convincing themselves of things that are tangential or opposite to the task

Funny that, when given too much brainpower, AIs manifest ADHD symptoms…

Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second

#97
post #55

Earlier quoted context omitted.

Well MemoryX compared to H100 HBM3 the key details are that MemoryX has lower latency, but also far lower bandwidth. However the memory on Cerebras is scales a lot more over NVidia. You need a cluster of H100's to create a model, as only way to scale the memory, Cerbras is more suited to that aspect, Nvidia do their scaling in tooling, with Cerbras doing theirs in design via there silicon approach. That's my take on…

No way an offchip HBM has same or better bandwidth then onchip

> MemoryX has lower latency, but also far lower bandwidth

Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second

#99

If this is the full fp16 quant, you'd need 2TB of memory to use with the full 131k context. With 44GB of SRAM per Cerebras chip, you'd need 45 chips chained together. $3m per chip. $135m total to run this. For comparison, you can buy a DGX B200 with 8x B200 Blackwell chips and 1.4TB of memory for around $500k. Two systems would give you 2.8TB memory which is enough for this. So $1m vs $135m to run this model. It's no…

Our chips don't cost $3M. I'm not sure where you got that number but its wildly incorrect.

Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second

#100
post #78

Earlier quoted context omitted.

Never will. But then, same for humans yes?

>But then, same for humans yes? And? Whats your point? This is a computer. Humans make errors doing arithmetic, therefore should we not expect computers to be able to reliably perform arithmetic? No. Silly retort and a common reply from people who are suitably wowed by the current generation of AI.

This is incredibly dumb.
Post reply on HN