Live data from Hacker News

Cerebras CS-4

cerebras.ai

181–190 of 281 posts

Re: Cerebras CS-4

#181

OpenAI needs to immediately move to acquire Cerebras. Nvidia's extreme margin is the opportunity for OpenAI's cost reduction. Buying Cerebras would pay for itself and they should take all of its future production (after filling required contracts). Right now China's models have no silicon moat. Cerebras as a drastic speed-up / cost-reduction potential, can assist in building a competitive moat. And every time a Cereb…

[deleted]

Re: Cerebras CS-4

#182

Earlier quoted context omitted.

I guess we know why there's a fair bit of investment money going into small modular nuclear reactor startups now.

Actually that's mostly just the military funding those, with a few of them having data center partnerships so they can shield themselves from the criticism of what they really are: military contractors.

I heard a great deal of noise from early 2002 to the present date that the US military had a high interest in small portable nuclear reactors for large bases in Iraq and Afghanistan. And particularly around the peak period of troops on the ground in AF and IQ. And indeed a place like Bagram or Kandahar used a shitton of diesel to run generators. But nothing ever came to fruition to actually implement it, has something changed now that they actually consider it worth doing?

Re: Cerebras CS-4

#183
God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing

Guess they don't care about regular devs atm and are focused only on hardware sales.

Re: Cerebras CS-4

#184

Earlier quoted context omitted.

(Where did you see that?) This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use? I think the frontier providers keep the size of their models carefully hidden.

pretty sure 10 trillion parameters is now the norm among closed ai labs, given that nvidia also references the same 10 trillion number for their nvl72 racks

Pretty bad efficiency then unless that only applies to Fable class but even then - Kimi K3 is around 3T and does similarly well in most benchmarks.

Re: Cerebras CS-4

#185

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

Cerebras capacity was pretty much entirely bought out at some point. We needed it and couldn't get it.

Re: Cerebras CS-4

#186
post #113

Earlier quoted context omitted.

or the latest qwen3.8 27B doing so well at ~1/100 the size of K3

What about general knowledge you can get out of it before hallucinations start?

Qwen 3.8 27B beats Opus, Fable and GPT 5.6 by a comfortable margin on the AA-Omniscience Hallucination Rate benchmark.

Re: Cerebras CS-4

#187

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

> GPT-OSS 120B which is nigh useless nowadays:

I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent.

I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

Re: Cerebras CS-4

#189
post #3

Earlier quoted context omitted.

This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new d…

By the way, this is the same argument that Michael Burry used to short Nvidia. He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0] The logic is fundamentally flawed in my opinion. Let'…

1. The same isn’t necessarily true of the rest of the hardware stack which may be reused between accelerator generations.

2. You’re missing the “New Nvidia chips 10x B200, compute requirement grows less than 10*software improvements YoY -> buy less Nvidia.” Valuations are based on forward projections (>1T annual for NVDA) which can be revised down leading to a drop in valuation.

> If Amazon doesn't buy but Microsoft does

The big 3 all have their own proprietary accelerators. Meta is buying TPUs as well for now.

I would bet Nvidia’s major customers in 2 years are neoclouds and it seems that Jensen is making the same bet.

Re: Cerebras CS-4

#190

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

I don’t think they will until they change the architecture.

They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.

Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.

Post reply on HN