Earlier quoted context omitted.
I thought your numbers must be wrong. So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum. So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.
For reference, Google serves >3 quadrillion/mo. [1] https://x.com/ren_stocks/status/2056946641815396718?s=20
Cerebras CS-4
231–240 of 281 posts
Re: Cerebras CS-4
#232God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…
Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them? OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.
Because they already have / had an okay coding subscription product for a bit and it gives them visibility and mindshare (in regards to their hardware, even if they don't compete with other providers that much). They could do what Kimi did - make a good subscription with good models, once you get enough customers to get some good PR and such, pause the signups so you don't have to spend more on running the service than you want/can. Do enough of that and people will talk about your offerings organically, make yourselves known to even devs as "That one company with their own hardware and the super fast subscription." experiencing which would do more than any marketing.
Re: Cerebras CS-4
#233God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…
Why they should go with Chinese models if they have a line up of gpt models and a very good partnership with someone who lives in the same jurisdiction and not in the country that convinces their citizen that it’s a good idea to go on war with western world ? Just curious ?
Because they generated some buzz and are near-SOTA and would be a great benchmark for a PoC subscription that doesn't necessarily aim to compete with other vendors at a similar scale (since their main business is the hardware). Mistral is conceptually cool but is lagging behind. I guess Muse Spark and Laguna would also be okay, just not as recognizable. Meanwhile both Kimi K3 and GLM 5.3 are near-SOTA in performance and considerable in size, a great choice for proving the platform!
As for the 2nd part of your question - that wasn't a relevant concern or consideration here, unless the models would be tainted to a degree to prevent them from having a good coding subscription that gets more developer mindshare towards what their chips can achieve and generate some good PR.
Re: Cerebras CS-4
#234Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…
Re: Cerebras CS-4
#235God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…
Re: Cerebras CS-4
#236Earlier quoted context omitted.
Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them? OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.
> Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them? Because they already have / had an okay coding subscription product for a bit and it gives them visibility and mindshare (in regards to their hardware, even if they don't compete with other providers tha…
Cerebras is a B2B hardware company. It feels like a distraction: think of the opportunity cost, and resources/headcount not working on other things that would drive more impact.
Should NVIDIA do a coding subscription too? I'm sure they can make money off it, but I think it would be -EV.
Re: Cerebras CS-4
#237God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…
>Guess they don't care about regular devs atm and are focused only on hardware sales. Why would they want to target regular devs right now? If they sold to regular devs instead of enterprises, the complaint wouldn't be about model choice, it'd be about how expensive it.
Re: Cerebras CS-4
#238Earlier quoted context omitted.
needs an sla that says power will never ever ever go out or else you will have a useless shattered plate of silicon.
Huh, why would it shatter if the power goes out?
there is an extensive and complicated cooling system that permeates the wafer. some of the cores are completely turned off because they fail qc (see the tsmc logo) if all of their neighbors have been going at full bore the thermal differential can cause stress fractures if the cooling system suddenly fails.
Re: Cerebras CS-4
#239Earlier quoted context omitted.
And advanced geothermal. Fervo Energy let's us get energy that's not based on burning fossil fuels but is, instead, able to produce energy from the ground.
Remove energy from the ground. I wonder what the consequences may be once we are cooling the underground at several MW/h.
Re: Cerebras CS-4
#240God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…
> Guess they don't care about regular devs atm and are focused only on hardware sales
They aren't trying to make a few bucks off tokenmaxxers. They're trying to be the underpinning of compute for all AI. They're going to beat Nvidia.