Live data from Hacker News

Cerebras CS-4

cerebras.ai

191–200 of 281 posts

Re: Cerebras CS-4

#191

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

> GPT-OSS 120B which is nigh useless nowadays: I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent. I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

> I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs and you have to fix a non-insignificant amount of it all manually: https://blog.kronis.dev/blog/i-blew-through-24-million-token...

Admittedly that post was before agentic development truly took off and that 3k EUR figure when paying per API tokens would nowadays be closer to like 6k EUR for the volume of work I do, but still.

It's the same how Qwen 2.5 was pretty problematic for anything remotely serious, same with Qwen 3 Coder Next (80B), and at least the most recent versions are getting better but still not quite good enough in real world use cases outside of benchmarks. They've come a long way, regardless!

Re: Cerebras CS-4

#192
post #40

Earlier quoted context omitted.

It's rumored fable is around that 10T number

If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.

And GLM is only 0.7T!

But these labs distill off the larger models. Both officially at the labs with the big ones, and unofficially. We need the giant models to get the smaller models.

Re: Cerebras CS-4

#194

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

> GPT-OSS 120B which is nigh useless nowadays: I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent. I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

As MoE with 5B active parameters it's pretty fast. But you still need a lot of vRAM, or have to run small quantitations. Qwen models just gave you more bang for your buck, and the gap became worse with every qwen release

Re: Cerebras CS-4

#195

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

I use Cerebras via OpenRouter. It’s every bit as fast and reliable for my needs as claimed. I suspect the reason is that they can either be making peanuts selling inference to plebs like me via OpenRouter, or making bank selling the more expensive models to businesses directly. In short: I would be very surprised if they have die capacity, and are at this point maximising revenue per chip.

Re: Cerebras CS-4

#196

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

Probably the same reason why there are more people who takes buses, subways, trains than drive Ferraris.

I love how instead of comparing a Ferrari (fast and expensive) to some average car (not fast, not expensive) to make your point..you went for public transport where your comparison cracks from multiple angles.

Re: Cerebras CS-4

#197

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

> GPT-OSS 120B which is nigh useless nowadays: I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent. I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

gpt-oss-120b is absolutely unusable over Cerebras. It fails to call tools half the time and just continues to think about what tool it'll call repeatedly. Like it says it'll call a tool and then it doesn't, and then it says it'll call the tool again and then it doesn't, and it just does that in a loop forever. It's awful. Also forgets to end the thinking block too. Even if the model itself was just-okay for its time, even at 1000t/s+ it's not worth it. And it's EXPENSIVE, like $5 per minute expensive

Re: Cerebras CS-4

#198
post #134

Earlier quoted context omitted.

I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months. I m sure my employer would too. Qwwen3.5 122b was released 6 months ago and is still best in class overall 100-140 B param model.

> I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months. .... :T Considering the 8B model uses 53 billion transistors, that's 6.625 transistors per parameter. https://taalas.com/products/ Assuming they can get it down to 3 (somehow), that's still 300 transistors, or 5.565 RX 9070s. https://www.techpowerup.com/gpu-specs/radeon-rx-9070.c4250 You're looking at…

In Taalas HC2 a chip embeds 20b parameters, and the declared idea is linking the chips. A card with two of them chips and you can already have a dense Qwen at staggering speeds.

Re: Cerebras CS-4

#199

Earlier quoted context omitted.

Rumors and allegations aren't worth much. Why don't they just tell us mere mortals?

Because that would reveal their edge to investors, or the lack thereof. If Fable turns out to be a 10T or 20T model, there is little to boast vs Kimi at 3T. But the opposite is true: if Fable were to be e.g. a 500B model, that would show how far ahead they are from the open models. This isn't likely to be the case ...

I guess there's also the economic aspect. It would make it much easier for competitors to figure out your costs and margins if they know the model parameter sizes you operate.

Re: Cerebras CS-4

#200

Earlier quoted context omitted.

> The absolute worst market time to etch a model to a chip is right now Slightly disagree. It really depends on the price-point at which they can do that etching. ~1k usd / ~30B model in a hdd-sized case that fits on your desk? I'd buy one right now, even knowing that I'm "stuck" with whatever model of the day is.

Time to market also matters a ton. If they can start shipping chips <1 month after the weights drop that's much more compelling than if it's a 6+ month development pipeline.

In the case of Taalas, the pipeline was said to be 2 months:

> From the moment a previously unseen model is received, it can be realized in hardware in only two months ( https://taalas.com/the-path-to-ubiquitous-ai/ )

Post reply on HN