God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…
Cerebras CS-4
271–280 of 281 posts
Re: Cerebras CS-4
#272Earlier quoted context omitted.
Why would they? What the upside, for them?
Indeed, what is the upside of transparency?
For a charity? Assuring donors that funds are being managed appropriately.
For Anthropic? No benefit at all.
Turns out "transparency" is like "weight" or "velocity" in that it has no intrinsic value, and can be positive or negative depending on context.
Re: Cerebras CS-4
#273Earlier quoted context omitted.
I just want to but hardware so I can run a model at home that is fast. I don't see myself installing a server that burns almost two hundred kilowatts but maybe a card which runs a 27B Qwen...
At 250w when it's working (I understand), and it works for tiny amounts of time per query...
Re: Cerebras CS-4
#274Earlier quoted context omitted.
In Taalas HC2 a chip embeds 20b parameters, and the declared idea is linking the chips. A card with two of them chips and you can already have a dense Qwen at staggering speeds.
Chiplet-style layouts could cut the etching quality requirements per chip down, but it still can't avoid the base cost for manufacturing silicon.
What costs are you talking about and why would them be a problem?
If it is the price: «Kharya says it costs 100x as much to train a model then to get a customize HC chip in reasonable volumes from Taalas» ( https://www.nextplatform.com/compute/2026/02/19/taalas-etche... ).
That is thinking about an LLM (logical) producer and server. For mass production, the costs go down. And in the case of a ~100b model as the poster mentioned, they would be just single PCI cards with 5 or 6 HC2 chips: doable and practical.
I see more potential problems in the positioning of the SRAM - but not a real problem given that excellent team.
To get a proper idea of the costs the architecture of the HC2 will have to be clearer.
Re: Cerebras CS-4
#275Earlier quoted context omitted.
Coding subs are good when they promote usage and adoption of your models in enterprises at API rates. Cerebras is a B2B hardware company. It feels like a distraction: think of the opportunity cost, and resources/headcount not working on other things that would drive more impact. Should NVIDIA do a coding subscription too? I'm sure they can make money off it, but I think it would be -EV.
> Should NVIDIA do a coding subscription too? Yes, obviously! Well maybe not a subscription but definitely an inference service. https://build.nvidia.com/ https://resources.nvidia.com/en-us-inference-infrastructure/... https://www.nvidia.com/en-us/data-center/dgx-cloud-lepton/ In their case not to gain mindshare or money or whatever, they're already a market leader, but to run something that validates the use case of…
Re: Cerebras CS-4
#276God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…
Re: Cerebras CS-4
#277Earlier quoted context omitted.
But power hungry. In that 5+ year timeline, the compute per watt could change by three orders of magnitude. GPUs are to LLMs what CPUs are to gaming — not a good fit.
A cursory estimate courtesy of ChatGPT suggests that there is a grand total of one order of magnitude or less of power efficiency improvement available compared to current Blackwell if the entire system’s power consumption outside the ALUs went all the way to zero. If you want three orders of magnitude improvement, you probably need to find two of those orders of magnitude somewhere else: process improvements, differ…
Re: Cerebras CS-4
#278Earlier quoted context omitted.
Why they should go with Chinese models if they have a line up of gpt models and a very good partnership with someone who lives in the same jurisdiction and not in the country that convinces their citizen that it’s a good idea to go on war with western world ? Just curious ?
> Why they should go with Chinese models Because they generated some buzz and are near-SOTA and would be a great benchmark for a PoC subscription that doesn't necessarily aim to compete with other vendors at a similar scale (since their main business is the hardware). Mistral is conceptually cool but is lagging behind. I guess Muse Spark and Laguna would also be okay, just not as recognizable. Meanwhile both Kimi K3…
You pick the vendor who is ahead, create an agreement with them to get access ahead of public release , and bake that model into hardware , because that’s how you make money .
Re: Cerebras CS-4
#279Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…
This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new d…
Why? Make your case.
Re: Cerebras CS-4
#280God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…