Cerebras is truly one of the maddest technical accomplishments that Silicon Valley has produced in the last decade or so. I met Andy seven or eight years ago and I thought they must have been smoking something - a dinner plate sized chip with six tons of clamping force? They made it real, and in retrospect what they did was incredibly prescient
The concept is super cool but does anyone actually use them instead of just buying Nvidia?
Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
91–100 of 160 posts
Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
#92If this is the full fp16 quant, you'd need 2TB of memory to use with the full 131k context. With 44GB of SRAM per Cerebras chip, you'd need 45 chips chained together. $3m per chip. $135m total to run this. For comparison, you can buy a DGX B200 with 8x B200 Blackwell chips and 1.4TB of memory for around $500k. Two systems would give you 2.8TB memory which is enough for this. So $1m vs $135m to run this model. It's no…
I'd think that HFT is already mature and doesn't really benefit from this type of model.
Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
#93Insane that Cerebras succeeded where everyone else failed for 5 decades.
Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
#94If this is the full fp16 quant, you'd need 2TB of memory to use with the full 131k context. With 44GB of SRAM per Cerebras chip, you'd need 45 chips chained together. $3m per chip. $135m total to run this. For comparison, you can buy a DGX B200 with 8x B200 Blackwell chips and 1.4TB of memory for around $500k. Two systems would give you 2.8TB memory which is enough for this. So $1m vs $135m to run this model. It's no…
>Maybe hedge funds or some sort of financial markets? I'd think that HFT is already mature and doesn't really benefit from this type of model.
Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
#95Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
#96With this kind of speed you could build a large thinking stage into every response. What kind of improvement could you expect in benchmarks from having say 1000 tokens of thinking for every response?
Thinking can also make the responses worse; AIs don't "overthink", instead they start throwing away constraints and convincing themselves of things that are tangential or opposite to the task. I've often observed thinking/reasoning to cause models to completely disregard important constraints, because they essentially can act as conversational turns.
Funny that, when given too much brainpower, AIs manifest ADHD symptoms…
Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
#97Earlier quoted context omitted.
Well MemoryX compared to H100 HBM3 the key details are that MemoryX has lower latency, but also far lower bandwidth. However the memory on Cerebras is scales a lot more over NVidia. You need a cluster of H100's to create a model, as only way to scale the memory, Cerbras is more suited to that aspect, Nvidia do their scaling in tooling, with Cerbras doing theirs in design via there silicon approach. That's my take on…
No way an offchip HBM has same or better bandwidth then onchip
Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
#98If they do the same for the coding model they will have a killer product.
Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
#99If this is the full fp16 quant, you'd need 2TB of memory to use with the full 131k context. With 44GB of SRAM per Cerebras chip, you'd need 45 chips chained together. $3m per chip. $135m total to run this. For comparison, you can buy a DGX B200 with 8x B200 Blackwell chips and 1.4TB of memory for around $500k. Two systems would give you 2.8TB memory which is enough for this. So $1m vs $135m to run this model. It's no…
Re: Cerebras launches Qwen3-235B, achieving 1.5k tokens per second
#100Earlier quoted context omitted.
Never will. But then, same for humans yes?
>But then, same for humans yes? And? Whats your point? This is a computer. Humans make errors doing arithmetic, therefore should we not expect computers to be able to reliably perform arithmetic? No. Silly retort and a common reply from people who are suitably wowed by the current generation of AI.