Live data from Hacker News

Cerebras CS-4

cerebras.ai

231–240 of 281 posts

Re: Cerebras CS-4

#231

Earlier quoted context omitted.

I thought your numbers must be wrong. So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum. So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.

For reference, Google serves >3 quadrillion/mo. [1] https://x.com/ren_stocks/status/2056946641815396718?s=20

Which is only 10x more than Open Router by the way.

Re: Cerebras CS-4

#232
post #214

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them? OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.

> Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them?

Because they already have / had an okay coding subscription product for a bit and it gives them visibility and mindshare (in regards to their hardware, even if they don't compete with other providers that much). They could do what Kimi did - make a good subscription with good models, once you get enough customers to get some good PR and such, pause the signups so you don't have to spend more on running the service than you want/can. Do enough of that and people will talk about your offerings organically, make yourselves known to even devs as "That one company with their own hardware and the super fast subscription." experiencing which would do more than any marketing.

Re: Cerebras CS-4

#233
post #226

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

Why they should go with Chinese models if they have a line up of gpt models and a very good partnership with someone who lives in the same jurisdiction and not in the country that convinces their citizen that it’s a good idea to go on war with western world ? Just curious ?

> Why they should go with Chinese models

Because they generated some buzz and are near-SOTA and would be a great benchmark for a PoC subscription that doesn't necessarily aim to compete with other vendors at a similar scale (since their main business is the hardware). Mistral is conceptually cool but is lagging behind. I guess Muse Spark and Laguna would also be okay, just not as recognizable. Meanwhile both Kimi K3 and GLM 5.3 are near-SOTA in performance and considerable in size, a great choice for proving the platform!

As for the 2nd part of your question - that wasn't a relevant concern or consideration here, unless the models would be tainted to a degree to prevent them from having a good coding subscription that gets more developer mindshare towards what their chips can achieve and generate some good PR.

Re: Cerebras CS-4

#234

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

What LLM-specific hardware improvements should one expect? Seems to me that LLM inference is simple architecturally (matmul et al) so most scaling in hardware should come from general improvements (memory BW, packaging, interconnect, power).

Re: Cerebras CS-4

#235

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

GLM 4.7 is gone (at least for us), with no suitable replacement from Cerebras. I think all they care about now is hardware and OpenAI hosting.

Re: Cerebras CS-4

#236
post #214

Earlier quoted context omitted.

Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them? OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.

> Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them? Because they already have / had an okay coding subscription product for a bit and it gives them visibility and mindshare (in regards to their hardware, even if they don't compete with other providers tha…

Coding subs are good when they promote usage and adoption of your models in enterprises at API rates.

Cerebras is a B2B hardware company. It feels like a distraction: think of the opportunity cost, and resources/headcount not working on other things that would drive more impact.

Should NVIDIA do a coding subscription too? I'm sure they can make money off it, but I think it would be -EV.

Re: Cerebras CS-4

#237

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

>Guess they don't care about regular devs atm and are focused only on hardware sales. Why would they want to target regular devs right now? If they sold to regular devs instead of enterprises, the complaint wouldn't be about model choice, it'd be about how expensive it.

Seriously, even as well paid as many devs are these are not machines that are affordable for personal use. Their market is people slapping down millions on frontier model training.

Re: Cerebras CS-4

#238
post #116

Earlier quoted context omitted.

needs an sla that says power will never ever ever go out or else you will have a useless shattered plate of silicon.

Huh, why would it shatter if the power goes out?

by my understanding:

there is an extensive and complicated cooling system that permeates the wafer. some of the cores are completely turned off because they fail qc (see the tsmc logo) if all of their neighbors have been going at full bore the thermal differential can cause stress fractures if the cooling system suddenly fails.

Re: Cerebras CS-4

#239

Earlier quoted context omitted.

And advanced geothermal. Fervo Energy let's us get energy that's not based on burning fossil fuels but is, instead, able to produce energy from the ground.

Remove energy from the ground. I wonder what the consequences may be once we are cooling the underground at several MW/h.

Earth will be swallowed by the sun before humans could make a dent in the earths core temperature.

Re: Cerebras CS-4

#240

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

Cerebras is the fastest provider by far on OpenRouter, and gpt-oss-120b is still very useful. They have backed up their claims very well.

> Guess they don't care about regular devs atm and are focused only on hardware sales

They aren't trying to make a few bucks off tokenmaxxers. They're trying to be the underpinning of compute for all AI. They're going to beat Nvidia.

Post reply on HN