Cerebras CS-4
211–220 of 281 posts
Re: Cerebras CS-4
#212Earlier quoted context omitted.
I guess we know why there's a fair bit of investment money going into small modular nuclear reactor startups now.
And advanced geothermal. Fervo Energy let's us get energy that's not based on burning fossil fuels but is, instead, able to produce energy from the ground.
Re: Cerebras CS-4
#213God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…
I don’t think they will until they change the architecture. They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks. Edit: I don’t know if they actually have a proper cache. This cou…
They do not seem to discount cached input for the self-serve Developer tier. Maybe they do for enterprise rate cards?
https://inference-docs.cerebras.ai/capabilities/prompt-cachi...
Re: Cerebras CS-4
#214God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing Guess they don't care about regular devs…
OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.
Re: Cerebras CS-4
#215I guess you'd need a DOZEN(s) of these to host a large model with long context KV caches?
Re: Cerebras CS-4
#216Earlier quoted context omitted.
The fact that Musk claims Opus is 5T to justify why Grok is far behind should be taken with a massive grain of salt given he's a recidivist mythomaniac. Honestly if Opus is 5T parameters while being matched by the biggest open models that are at least twice smaller, it would mean that the US is already behind China in the AI race, despite a significant edge in compute.
The open models don't really match Opus. For example I regularly do Fable+Opus agentic coding runs over 24 hours without intervention. I think I've had GLM do a run that was a few hours. That's the closest I've had an open model come on that kind of work.
Re: Cerebras CS-4
#217Re: Cerebras CS-4
#218I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.
Didn't they say that they can support bigger models now?
And there is, of course, the educated guesses about what WSE-4 will be, one being adding a LOT of stacked SRAM or DRAM to tip the balance towards memory (which could also be done by having a few different tile designs with various configurations of compute and memory capacity). I am curious about which way they'll go.
Re: Cerebras CS-4
#219Earlier quoted context omitted.
You just proved that AI cannot currently do that
It can definitely create a software stack for you if you hold it right, but the software stack supported by a trillion dollar company with decades of expertise, that also uses AI to improve its stack is probably gonna be better.
Re: Cerebras CS-4
#220Earlier quoted context omitted.
What do they do with old ones? Their hardware physically can't run other models right?
The Cerebras hardware is not locked to specific models / model families. Taalas is the company that's etching models into their silicon, locking it to that model forever.