If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
Cerebras CS-4
61–70 of 281 posts
Re: Cerebras CS-4
#62Earlier quoted context omitted.
162 kW
I guess we know why there's a fair bit of investment money going into small modular nuclear reactor startups now.
Re: Cerebras CS-4
#63Earlier quoted context omitted.
TPUs and ASICs run in data centers too. Your argument only holds true if there's some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.
Everything has some satisfied limit to demand, often depending on the price. If you assume there will never be any satisfied limit to demand for inference at any price you can justify any investment.
Re: Cerebras CS-4
#64Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…
I'm honestly baffled they were not acquired by somebody else (sorry AMD).
Re: Cerebras CS-4
#65Earlier quoted context omitted.
It's rumored fable is around that 10T number
If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.
There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today
Re: Cerebras CS-4
#66Earlier quoted context omitted.
They do offer API services to individual users... though with a set of models that makes it unlikely that you want to use it. They are promising Qwen 3.8 27B any day now though*. if you have the money as an "individual user" to purchase one of their racks... save your money and retire. * Actually they sent out an email claiming they already have it, but I don't seem to have access, they're promising to release it to…
> save your money and retire. Now that this hypothetical person has retired, what are they gonna do all day? Just sit on the beach and drink Mai Tais? If that's what they wanna do, sure, but nerds gonna nerd, and if I had that kind of money to retire on, I'd totally buy some ridiculously expensive AI box for fun.
Re: Cerebras CS-4
#67Re: Cerebras CS-4
#68Earlier quoted context omitted.
If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.
Not necessarily, there could be diminishing returns on mere parameters count . There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today
Re: Cerebras CS-4
#69Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…
This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new d…
A rough analogy would be if the first generation of ISP's spent billions on dial-up exchanges, when fibre could be invented next year.
Re: Cerebras CS-4
#70Earlier quoted context omitted.
Not necessarily, there could be diminishing returns on mere parameters count . There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today
That's precisely what he is saying, there is diminishing returns (or optimization left on the table).