Live data from Hacker News

Cerebras CS-4

cerebras.ai

241–250 of 281 posts

Re: Cerebras CS-4

#241

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

Because they aren't selling inference, they're selling hardware. The only reason they sell any tokens on OpenRouter is so they get on the benchmark that shows them as the fastest provider. It's free advertising.

Re: Cerebras CS-4

#242
post #40

Earlier quoted context omitted.

It's rumored fable is around that 10T number

Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.

Yeah, that's insane, you'd need to have many billions of dollars and buy up a huge chunk of the worlds memory supply to do that /s

Re: Cerebras CS-4

#243

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

What LLM-specific hardware improvements should one expect? Seems to me that LLM inference is simple architecturally (matmul et al) so most scaling in hardware should come from general improvements (memory BW, packaging, interconnect, power).

What you describe is basically Cerebras case, at the bottom it's just a really big die (about x28 an NVIDIA GB200) with a lot of work to reduce memory latency and improve throughput. What it's actually amazing is how can they make a chip so big and still have a decent yield to be commercially viable.

Re: Cerebras CS-4

#244
post #59

Earlier quoted context omitted.

It's mandatory liquid cooling, so it's meant to be attached to a specialized liquid cooling loop that gets the heat outside the building. This is far beyond the practical maximums of like 10 to 15kW per 44U cabinet front to rear air cooling for 'regular' rackmount server stuff.

Indeed. You need 45 to 60 liters per second of cooling water flowing over a Cerebras wafer every minute to keep it under 90C. And that’s assuming the water leaves at 90C… More realistically, you need much more cooling water.

CS-3 used 100 liters of water with a cold plate (https://www.brownstoneresearch.com/bleeding-edge/ai-infrastr...). But you can also use refrigerants with a cold plate (https://eng.umd.edu/engineering-ai-public-good/cooling-data-...) or dielectric fluid in total immersion. Until they find a more power-efficient design, my guess is total immersion will come back in style.

Re: Cerebras CS-4

#245
post #67
post #56

Earlier quoted context omitted.

I’ll get that 250kW home power service dropped in next week!

For now you can rig an adapter to your nearest DC EV charging station, but make sure it's near a body of water for the cooling.

Hot water for the whole neighbourhood!

Re: Cerebras CS-4

#246
post #214

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…

Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them? OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.

They do have Cerebras Code

https://www.cerebras.ai/code

But it's not fully open to just anyone, I wasted time signing up to find out that I couldn't even sign up for it to test it out.

Re: Cerebras CS-4

#247
post #213

Earlier quoted context omitted.

I don’t think they will until they change the architecture. They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks. Edit: I don’t know if they actually have a proper cache. This cou…

Cerebras supports prompt caching and has a doc about it. A fairly standard automatic prefix-based implementation with 5min expiry. They do not seem to discount cached input for the self-serve Developer tier. Maybe they do for enterprise rate cards? https://inference-docs.cerebras.ai/capabilities/prompt-cachi...

ah, that makes it feasible! Okay, glad it's not technical limit. They should fix the pricing...

Re: Cerebras CS-4

#248

Earlier quoted context omitted.

Remove energy from the ground. I wonder what the consequences may be once we are cooling the underground at several MW/h.

Earth will be swallowed by the sun before humans could make a dent in the earths core temperature.

The core is obviously not in question here, but at scale and long term it could have an impact on the water bed, vegetation, underground ecosystems, soil stability, etc.

Re: Cerebras CS-4

#249

can these vibe coded sites please set a max width and overflow so their sites work fine on mobile

Good news, future models will have your comment in their training set, making them slightly more likely to fix that problem!

What a time to be alive

Re: Cerebras CS-4

#250
post #103

Earlier quoted context omitted.

Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.

Fable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T. Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai... That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx Both Grok and Bytedance are training 10T models.

If Fable is seriously around 10T and Kimi K3 sidles up to it at 2.4T

That would be extremely surprising and a massive blunder by Anthropic in model design architecture ... which I highly doubt to be the case.

Post reply on HN