If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
Cerebras CS-4
241–250 of 281 posts
Re: Cerebras CS-4
#242Earlier quoted context omitted.
It's rumored fable is around that 10T number
Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.
Re: Cerebras CS-4
#243Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…
What LLM-specific hardware improvements should one expect? Seems to me that LLM inference is simple architecturally (matmul et al) so most scaling in hardware should come from general improvements (memory BW, packaging, interconnect, power).
Re: Cerebras CS-4
#244Earlier quoted context omitted.
It's mandatory liquid cooling, so it's meant to be attached to a specialized liquid cooling loop that gets the heat outside the building. This is far beyond the practical maximums of like 10 to 15kW per 44U cabinet front to rear air cooling for 'regular' rackmount server stuff.
Indeed. You need 45 to 60 liters per second of cooling water flowing over a Cerebras wafer every minute to keep it under 90C. And that’s assuming the water leaves at 90C… More realistically, you need much more cooling water.
Re: Cerebras CS-4
#245Re: Cerebras CS-4
#246God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…
Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them? OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.
But it's not fully open to just anyone, I wasted time signing up to find out that I couldn't even sign up for it to test it out.
Re: Cerebras CS-4
#247Earlier quoted context omitted.
I don’t think they will until they change the architecture. They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks. Edit: I don’t know if they actually have a proper cache. This cou…
Cerebras supports prompt caching and has a doc about it. A fairly standard automatic prefix-based implementation with 5min expiry. They do not seem to discount cached input for the self-serve Developer tier. Maybe they do for enterprise rate cards? https://inference-docs.cerebras.ai/capabilities/prompt-cachi...
Re: Cerebras CS-4
#248Earlier quoted context omitted.
Remove energy from the ground. I wonder what the consequences may be once we are cooling the underground at several MW/h.
Earth will be swallowed by the sun before humans could make a dent in the earths core temperature.
Re: Cerebras CS-4
#249Re: Cerebras CS-4
#250Earlier quoted context omitted.
Fable is most definitely nowhere near 10T. The cost to train and infer that would be insane, even by today's standards.
Fable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T. Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai... That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx Both Grok and Bytedance are training 10T models.
That would be extremely surprising and a massive blunder by Anthropic in model design architecture ... which I highly doubt to be the case.