Live data from Hacker News

Cerebras CS-4

cerebras.ai

271–280 of 281 posts

Re: Cerebras CS-4

#271

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

Wouldn't you want to keep this for internal development? Keep the single thread gains for yourself.

Re: Cerebras CS-4

#272
post #86

Earlier quoted context omitted.

Why would they? What the upside, for them?

Indeed, what is the upside of transparency?

For Linux? Assuring stakeholders that they can both control and observe development and how it works.

For a charity? Assuring donors that funds are being managed appropriately.

For Anthropic? No benefit at all.

Turns out "transparency" is like "weight" or "velocity" in that it has no intrinsic value, and can be positive or negative depending on context.

Re: Cerebras CS-4

#273
post #120

Earlier quoted context omitted.

I just want to but hardware so I can run a model at home that is fast. I don't see myself installing a server that burns almost two hundred kilowatts but maybe a card which runs a 27B Qwen...

At 250w when it's working (I understand), and it works for tiny amounts of time per query...

Corrige: a Taalas HC1 PCIe card consumes actually between ~200W and ~250W (a 10 blades server consumes ~2500W).

Re: Cerebras CS-4

#274

Earlier quoted context omitted.

In Taalas HC2 a chip embeds 20b parameters, and the declared idea is linking the chips. A card with two of them chips and you can already have a dense Qwen at staggering speeds.

Chiplet-style layouts could cut the etching quality requirements per chip down, but it still can't avoid the base cost for manufacturing silicon.

> it still can't avoid the base cost for manufacturing silicon

What costs are you talking about and why would them be a problem?

If it is the price: «Kharya says it costs 100x as much to train a model then to get a customize HC chip in reasonable volumes from Taalas» ( https://www.nextplatform.com/compute/2026/02/19/taalas-etche... ).

That is thinking about an LLM (logical) producer and server. For mass production, the costs go down. And in the case of a ~100b model as the poster mentioned, they would be just single PCI cards with 5 or 6 HC2 chips: doable and practical.

I see more potential problems in the positioning of the SRAM - but not a real problem given that excellent team.

To get a proper idea of the costs the architecture of the HC2 will have to be clearer.

Re: Cerebras CS-4

#275
post #236

Earlier quoted context omitted.

Coding subs are good when they promote usage and adoption of your models in enterprises at API rates. Cerebras is a B2B hardware company. It feels like a distraction: think of the opportunity cost, and resources/headcount not working on other things that would drive more impact. Should NVIDIA do a coding subscription too? I'm sure they can make money off it, but I think it would be -EV.

> Should NVIDIA do a coding subscription too? Yes, obviously! Well maybe not a subscription but definitely an inference service. https://build.nvidia.com/ https://resources.nvidia.com/en-us-inference-infrastructure/... https://www.nvidia.com/en-us/data-center/dgx-cloud-lepton/ In their case not to gain mindshare or money or whatever, they're already a market leader, but to run something that validates the use case of…

Ultimately the frontier labs are competitors of Nvidia. There is a fixed amount that the market will pay for tokens. If the frontier lab model premium collapses due open weights models, Nvidia can capture a greater share of aggregate spend.

Re: Cerebras CS-4

#276

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

Didn't OpenAI bought it?

Re: Cerebras CS-4

#277
post #77

Earlier quoted context omitted.

But power hungry. In that 5+ year timeline, the compute per watt could change by three orders of magnitude. GPUs are to LLMs what CPUs are to gaming — not a good fit.

A cursory estimate courtesy of ChatGPT suggests that there is a grand total of one order of magnitude or less of power efficiency improvement available compared to current Blackwell if the entire system’s power consumption outside the ALUs went all the way to zero. If you want three orders of magnitude improvement, you probably need to find two of those orders of magnitude somewhere else: process improvements, differ…

Oh, you could go analog rather than digital. Ever look at ALU design? Nothing efficient about it!

Re: Cerebras CS-4

#278
post #226

Earlier quoted context omitted.

Why they should go with Chinese models if they have a line up of gpt models and a very good partnership with someone who lives in the same jurisdiction and not in the country that convinces their citizen that it’s a good idea to go on war with western world ? Just curious ?

> Why they should go with Chinese models Because they generated some buzz and are near-SOTA and would be a great benchmark for a PoC subscription that doesn't necessarily aim to compete with other vendors at a similar scale (since their main business is the hardware). Mistral is conceptually cool but is lagging behind. I guess Muse Spark and Laguna would also be okay, just not as recognizable. Meanwhile both Kimi K3…

So they should go with near sota instead of SoTA due to hype ? In the Current market there are vendors who are ahead and Chinese quickly catch up.

You pick the vendor who is ahead, create an agreement with them to get access ahead of public release , and bake that model into hardware , because that’s how you make money .

Re: Cerebras CS-4

#279
post #3

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new d…

> We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear.

Why? Make your case.

Re: Cerebras CS-4

#280

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

Exactly, they need to push to regular consumers as well as businesses. This will create more pressure for adoption
Post reply on HN